Encoder, decoder, encoding method and decoding method employing frequency domain prediction of tonal signals with time-varying pitches
EFDJHP addresses the inefficiencies of constant fundamental frequency assumptions in audio coding by employing a linearly changing pitch model, achieving improved encoding and decoding efficiency for speech and music with time-varying pitches.
Patent Information
- Application Number
- PCT/EP2025/057886
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-22
- Filing Date
- 2025-03-21
- Publication Date
- 2025-09-25
AI Technical Summary
Existing audio coding technologies struggle with inefficiencies in long-term prediction due to assumptions of constant fundamental frequencies, particularly in scenarios with time-varying pitches, which are prevalent in speech and music, leading to smearing effects and reduced performance, especially at higher frequencies.
The Extended Frequency Domain Joint Harmonics Prediction (EFDJHP) algorithm assumes linearly changing pitches between frames, using a closed-form mathematical representation to estimate harmonic parameters through a linear equation system, allowing for improved encoding and decoding of audio signals with time-varying pitches.
EFDJHP enhances coding efficiency by accommodating pitch variations, offering significant bitrate savings and improved performance in encoding speech and music signals with frequent pitch changes, outperforming previous methods in listening tests and bitrate comparisons.
Smart Images

Figure EP2025057886_25092025_PF_FP_ABST
Abstract
Description
[0001] Encoder, Decoder, Encoding Method and Decoding Method employing Frequency Domain Prediction of Tonal Signals with Time-Varying Pitches Description The present invention relates audio signal encoding and decoding and, in particular, to an encoder, a decoder, an encoding method and a decoding method employing frequency domain prediction of tonal signals with time-varying pitches. Long Term Prediction is a technique that exploits the periodicity in a sequence for the prediction of current samples. It is an important audio coding tool to remove the inter- frame redundancy in tonal signals such as speech and music [1]. In transform audio coding, the audio signal is firstly decomposed into time-frequency tiles using filterbanks such as Modified Discrete Cosine Transform (MDCT). The transform itself serves to reduce the long term redundancy by representing the harmonic components as single spectral lines. However, the transform is not adequate to represent the harmonics when the harmonics overlap with each other. A Long-Term Prediction (LTP) unit can be added to a transform audio coding to enhance the coding efficiency. Most existing solutions to long term prediction in transform domain audio coding are based on a naive assumption of a constant fundamental frequency of the tonal signals [1]–[4], while in realistic scenarios, their fundamental frequency usually changes with time. That is mostly prominent in speech signals where the frequent pitch variation is essential in distinguishing words and meanings, as well as in carrying emotion. Depending on the performing techniques music signals such as singing voice and solo instrumental music can also contain pitch variations. In the modern median, mixed-content items are immense. Thus a unified speech and audio codec, where both contents are coded in the transform domain, without mode switching as in Unified Speech and Audio Coding (USAC), is needed. Improving the LTP efficiency in case of pitch variations can enhance greatly the performance of a unified speech and audio coder. When the fundamental frequency varies, it is shown as smearing in the frequency domain. That smearing effect is especially dominant at a higher frequency end since the higher harmonics have a larger frequency changing ratio. The limited frequency resolution and the pitch variations restrict the performance of current LTP algorithms [5]. In low-bitrates coding where a longer frame length is used for redundancy reduction, the larger pitch variation within the frame proposes a severe challenge for LTP. That problem is mitigated in [5] by dividing the analysis frames into sub-frames. Adaptive window switching was used in [6]. In [7], the analysis filterbank in transform audio coding is time-warped the pitch contour on a frame basis, to accommodate the pitch variation within a frame. In [8], time warping on the time domain LTP filters is used to accommodate pitch variations in polyphonic signals. The Frequency Domain Joint Harmonics Prediction (FDJHP) algorithm proposed in [1] views the tonal components in an MDCT frame as the superposition of harmonics. Assuming that the fundamental frequencies are known and do not change between successive frames, the amplitudes and phases of those harmonics can be estimated by least mean square solving of a linear equation system, using the quantized and reconstructed MDCT frames of the most previous frames. The next frame are predicted based on the phase progression of the harmonics and the estimated harmonic parameters. The FDJHP algorithm shows very good coding efficiency in terms of bitrates savings and performs well in a listening test. The object of the present invention is to provide improved concepts for audio signal encoding and audio signal decoding. The object of the present invention is solved by the subject-matter of the independent claims. Preferred embodiments are provided in the dependent claims. An encoder according to an embodiment is provided. The encoder is configured for encoding a current frame of an audio signal depending on one or more previous frames of the audio signal. The one or more previous frames precede the current frame. Each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal. Moreover, each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain, The encoder is configured to determine a change of a pitch or an estimation of a change of the pitch between a most previous frame of the one or more previous frames and a current frame. Furthermore, the encoder is configured to generate an encoding of the current frame depending on an estimation of two harmonic parameters for each of one or more harmonic components of the most previous frame, and depending on the change of the pitch or the estimation of the change of the pitch. Furthermore, a decoder according to an embodiment is provided. The decoder is configured for reconstructing a current frame of an audio signal. One or more previous frames of the audio signal precede the current frame. Each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal. Moreover, each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain. The decoder is to receive an encoding of the current frame. Moreover, the decoder is to determine an estimation of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames. Furthermore, the decoder is to reconstruct the current frame depending on the encoding of the current frame. The decoder is to reconstruct the current frame depending on the estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame. Moreover, the decoder is to reconstruct the current frame depending on a change of a pitch or an estimation of a change of the pitch between a most previous frame of the one or more previous frames and a current frame. Moreover, a system according to an embodiment is provided. The system comprises an encoder according to an embodiment and a decoder according to an embodiment. The decoder is configured to receive the encoding of the current frame from the encoder. Furthermore, a method for encoding a current frame of an audio signal depending on one or more previous frames of the audio signal according to an embodiment is provided. The one or more previous frames precede the current frame. wherein Each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal. Moreover, each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain. The method comprises: - Determining a change of a pitch or an estimation of a change of the pitch between a most previous frame of the one or more previous frames and a current frame. And: - Generating an encoding of the current frame depending on an estimation of two harmonic parameters for each of one or more harmonic components of the most previous frame, and depending on the change of the pitch or the estimation of the change of the pitch. Moreover, a method for reconstructing a current frame of an audio signal according to an embodiment is provided. One or more previous frames of the audio signal precede the current frame. Each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal. Each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain. The method comprises: - Receiving an encoding of the current frame. - Determining an estimation of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames. And: - Reconstructing the current frame depending on the encoding of the current frame. Reconstructing the current frame is conducted depending on the estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame. Moreover, reconstructing the current frame is conducted depending on a change of a pitch or an estimation of a change of the pitch between a most previous frame of the one or more previous frames and a current frame. For example, in a backward adaptive prediction scheme, the harmonic parameters may, e.g., be estimated at both the encoder and the decoder on a frame basis, and thus are not part of the encoding of a current frame and are not transmitted to the decoder. A prediction unit comprising the EFDJHP algorithm of an embodiment may, e.g., operate in a backward adaptive fashion, where reconstructed spectral coefficients may, e.g., be used in the harmonic estimation step of the proposed EFDJHP algorithm, both at the encoder and at the decoder. Thus the estimated harmonic parameters don’t need to be transmitted to the decoder. In the encoder, a refined fundamental frequency and an optimal gain factor may, e.g., be determined on a frame basis by minimizing the perceptual entropy of the quantized residual signal. The refined fundamental frequency and the optimal gain factor may, e.g., then be transmitted to the decoder, e.g., as part of the side information. Furthermore, a computer program according to an embodiment for implementing the one of the above-described methods when being executed on a computer or signal processor is provided. According to an embodiment, a codec with a frequency domain long term prediction unit can be applied to a broader range of input signals is provided. In comparison to a previous work where the pitch of the tonal part in the input audio signal is assumed to be constant between neighboring frames [1], in this present invention, the pitch is assumed to be changing linearly between neighboring frames. The previous frequency domain long term prediction algorithm described in [1] has also been published as a journal [2], and is denoted here as Frequency Domain Joint Harmonics Prediction (FDJHP). In this present invention, the improved frequency domain long term prediction algorithm is denoted as Extended Frequency Domain Joint Harmonics Prediction (EFDJHP). The codec with a frequency domain long term prediction unit comprises further following concepts: receive audio input signals, transform the input signal from time domain to frequency domain, apply a pitch tracking algorithm to the input signal to extract the pitch track, apply a psychoacoustic model to the input signal to extract quantization stepsizes, apply quantization to the residual spectra, and apply entropy coding to the quantized residual spectra. Embodiments extend a previously proposed FDJHP algorithm to cope with input signals with pitch variations. According to embodiments, the signal model used in a previously proposed Frequency Domain Joint Harmonics Prediction (FDJHP) algorithm, which does long term prediction for transform domain general audio coding, with the assumption of a constant fundamental frequency between adjacent frames, has been extended. This extended model according to embodiments assumes that the tonal input signals have time-varying pitches. According to an embodiment, a closed-form algorithm extension to predict tonal signals with linearly changing pitches is provided. Experiments show that this Extended Frequency Domain Joint Harmonics Prediction (EFDJHP) algorithm of an embodiment can accept pitch updates when the pitch variation is not linear, and still offers a significant improvement in coding gain for tonal signals with frequent pitch variations such as singing voices and speech. The Extended Frequency Domain Joint Harmonics Prediction (EFDJHP) of an embodiment extends the signal model and assumes that the instantaneous frequency of the tonal part may change linearly within a short time segment. In an embodiment, a closed-form mathematical representation of the MDCT frames is provided as a linear equation system, which can be solved once the fundamental frequencies are known. The algorithm extension of embodiments is shown to improve the prediction efficiency. In the following, embodiments of the present invention are described in more detail with reference to the figures, in which: Fig.1 illustrates a system according to an embodiment, comprising an encoder according to an embodiment and a decoder according to an embodiment. Fig.2 illustrates an improvement of coding gain using EFDJHP on sweep signals with linearly changing fundamental frequencies, wherein the MDCT frame length is 256. Fig.3(a)–3(c) illustrate a bitrate saving comparison between AACLTP. FDSHP, FDJHP and EFDJHP, wherein the prediction bandwidth has been selected to be 4 kHz. Fig. 1 illustrates an encoder 100 according to an embodiment. The encoder 100 is configured for encoding a current frame of an audio signal depending on one or more previous frames of the audio signal. The one or more previous frames precede the current frame. Each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal. Moreover, each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain, The encoder 100 is configured to determine a change of a pitch or an estimation of a change of the pitch between a most previous frame of the one or more previous frames and a current frame. Furthermore, the encoder 100 is configured to generate an encoding of the current frame depending on an estimation of two harmonic parameters for each of one or more harmonic components of the most previous frame, and depending on the change of the pitch or the estimation of the change of the pitch. Moreover, Fig.1 illustrates a decoder 200 according to an embodiment. The decoder 200 is configured for reconstructing a current frame of an audio signal. One or more previous frames of the audio signal precede the current frame. Each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal. Moreover, each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain. The decoder 200 is to receive an encoding of the current frame. Moreover, the decoder 200 is to determine an estimation of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames. Furthermore, the decoder 200 is to reconstruct the current frame depending on the encoding of the current frame. The decoder 200 is to reconstruct the current frame depending on the estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame. Moreover, the decoder 200 is to reconstruct the current frame depending on a change of a pitch or an estimation of a change of the pitch between a most previous frame of the one or more previous frames and a current frame. In the following, an encoder 100 according to particular embodiments is provided. An encoder 100 for encoding a current frame of an audio signal depending on one or more previous frames of the audio signal according to an embodiment is provided. The one or more previous frames precede the current frame, wherein each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal, wherein each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain. The most previous frame may, e.g., be most previous with respect to the current frame. The most previous frame may, e.g., be (referred to as) an immediately preceding frame. The immediately preceding frame may, e.g., immediately precede the current frame. The current frame comprises one or more harmonic components of the audio signal. Each of the one or more previous frames might comprise one or more harmonic components of the audio signal. The fundamental frequency of the one or more harmonic components in the current frame and the one or more previous frames are assumed known. The fundamental frequency may, e.g., be assumed changing linearly over the current frame and the one or more previous frames. To generate an encoding of the current frame, the encoder 100 is to determine an estimation of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames. Moreover, the encoder 100 is to determine the estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame using a first group of three or more of the plurality of spectral coefficients of the one or more previous frames of the audio signal. According to an embodiment, the encoder 100 may, e.g., be configured to determine a gain factor and a residual signal as the encoding of the current frame depending on a fundamental frequency of the one or more harmonic components of the current frame and the one or more previous frames and depending on the estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame. The encoder 100 may, e.g., be configured to generate the encoding of the current frame such that the encoding of the current frame comprises the gain factor and the residual signal. In an embodiment, the encoder 100 may, e.g., be configured to determine an estimation of the two harmonic parameters for each of the one or more harmonic components of the current frame depending on the estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame and depending on the fundamental frequency of the one or more harmonic components of the current frame and the one or more previous frames. The fundamental frequency may, e.g., be assumed changing linearly over the current frame and the one or more previous frames. According to an embodiment, the two harmonic parameters for each of the one or more harmonic components are a first parameter for a cosinus sub-component and a second parameter for a sinus sub-component for each of the one or more harmonic components. In an embodiment, the encoder 100 may, e.g., be configured to estimate the two harmonic parameters for each of the one or more harmonic components of the most previous frame by solving a linear equation system comprising at least three equations, wherein each of the at least three equations depends on a spectral coefficient of the first group of the three or more of the plurality of spectral coefficients of the one or more previous frames. According to an embodiment, the encoder 100 may, e.g., be configured to solve the linear equation system using a least mean squares algorithm. According to an embodiment, the linear equations using only the most previous frame are defined by wherein m – 1 is a frame index of the most previous frame, wherein Г is a set of subband indices of length Lm–1, in which the one or more harmonic components lie in the most previous frame, wherein p is a first vector comprising the two harmonic parameters for each of the one or more harmonic components of the most previous frame: wherein ahis a parameter for a cosinus sub-component for an h-th harmonic component of the most previous frame, wherein bhis a parameter for a sinus sub-component for the h-th harmonic component of the most previous frame, wherein H indicates a number of the harmonic components of the most previous frame, wherein Um–1comprises a number of third matrices or third vectors: wherein the third matrix or third vector for an h-th harmonic component of the most previous frame is defined by: 10 (4b) wherein is defined by:
[0002] 5 (5) wherein is defined by:
[0003] (6)
[0004] 10 wherein
[0005] 15 is the instantaneous angular frequency at time stamp of the h-th harmonic component, wherein is fundamental frequency of the one or more harmonic components at time stamp , wherein fs is a sampling frequency,
[0006] 20 wherein is a set of subband indices, in which the h-th harmonic component lie in the most previous frame, wherein
[0007] 25 is the modulation frequency in subbands defined by wherein
[0008] 30 (7) wherein wherein Bm–1,hrepresents the changing ratio of the angular frequency of the h-th harmonic component in the most previous frame, and for the special case of constant f 0, i.e., B = 0, leads to and equation (3) descends to the identical form of the Modified Discrete Cosine Transform (MDCT) domain representation in [1] for constant fundamental frequencies, wherein f (n) is the (symmetric) window function, wherein N depends on a length of a transform block for transforming the time-domain audio signal into the frequency domain or into the spectral domain. According to an embodiment, linear equations using more previous frames may be added to construct a solvable linear equation system, where the number of previous frames needed is denoted by S. According to an embodiment, when the fundamental frequency of the one or more harmonic components does not change linearly in the current frame and the one or more previous frames, an approximation of a linearly changing fundamental frequency of the one or more harmonic components can be applied in the current frame and the one or more previous frames, where the changing ratio of the angular frequency of the one or more harmonic components may be assumed to be the same in the current frame and the one or more previous frames, wherein the changing ratio of the angular frequency of the h-th harmonic component in the current frame and the one or more previous frames can be obtained by: with s = 1, …, S, wherein f 0mis the fundamental frequency of the one or more harmonic components at the center of the m-th frame. Furthermore, an approximation of the fundamental frequency of the one or more harmonic components in the one or more previous frames can be obtained from a linear extrapolation using the approximated changing ratio of the angular frequency and the fundamental frequencies of the one or more harmonic components in the current frame and the most previous frame: with s = 1, …, S. Alternatively, according to an embodiment, when the fundamental frequency of the one or more harmonic components does not change linearly in the current frame and the one or more previous frames, the fundamental frequency of the one or more harmonic components could be assumed piece-wisely linearly changing in two neighboring frames of the current frame and the one or more previous frames, wherein In an embodiment, the linear equation system is solvable according to: (11) wherein is a first vector comprising an estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame, wherein is a second vector comprising the first group of the three or more of the plurality of spectral coefficients of the one or more previous frames, wherein is a Moore-Penrose inverse matrix of U, wherein U comprises a number of third matrices or third vectors of the one or more previous frames. In an embodiment, the encoder 100 may, e.g., be to encode a fundamental frequency of harmonic components, a window function, the gain factor and the residual signal. According to an embodiment, the encoder 100 may, e.g., be configured to determine the number of the one or more harmonic components of the most previous frame and a fundamental frequency of the one or more harmonic components of the most previous frame before estimating the two harmonic parameters for each of the one or more harmonic components of the most previous frame using a first group of three or more of the plurality of spectral coefficients for each of the one or more previous frames of the audio signal. According to an embodiment, the encoder 100 may, e.g., be configured to predict the harmonic parameters of the one or more harmonic components in the current frame, denoted as: wherein chis a parameter for a cosinus sub-component for the h-th harmonic component of said one or more harmonic components of the current frame, wherein dhis a parameter for a sinus sub-component for the h-th harmonic component of said one or more harmonic components of the current frame: wherein In an embodiment, the encoder 100 may, e.g., be configured to determine a spectral prediction of one or more of the plurality of spectral coefficients of the current frame depending on the estimation of the two harmonic parameters for each of the one or more harmonic components of the current frame: wherein is a third vector comprising the spectral prediction of the three or more of the plurality of spectral coefficients of the current frame, wherein is a set of subband indices of length Lm, in which the one or more harmonic components lie in the current frame, wherein may be different from Г, when the fundamental frequency of the one or more harmonic components changes between the most previous frame and the current frame. wherein Umcomprises a number of third matrices or third vectors: wherein the third matrix or third vector for an h-th harmonic component of the current frame is defined by: wherein Ƴ2his defined by: is the instantaneous angular frequency at time stamp of the h-th harmonic component, wherein is the fundamental frequency of the one or more harmonic components at time stamp , wherein fs is a sampling frequency, wherein Гm,his a set of subband indices, in which the h-th harmonic component lie in the most previous frame, wherein is the modulation frequency in subbands defined by Гm,h, wherein wherein wherein Bm,hrepresents the changing ratio of the angular frequency of the h-th harmonic component in the current frame, and for the special case of constant f 0, i.e., B = 0, leads to and equation (15) descends to the identical form of the Modified Discrete Cosine Transform (MDCT) domain representation in [1] for constant fundamental frequencies, wherein f (n) is the (symmetric) window function, wherein N depends on a length of a transform block for transforming the time-domain audio signal into the frequency domain or into the spectral domain. The encoder 100 may, e.g., be configured to determine the residual signal and a gain factor depending on the plurality of spectral coefficients of the current frame in the frequency domain or in the transform domain and depending on the spectral prediction of the three or more of the plurality of spectral coefficients of the current frame; wherein the encoder 100 may, e.g., be configured to generate the encoding of the current frame such that the encoding of the current frame comprises the residual signal and the gain factor. According to an embodiment, the encoder 100 may, e.g., be configured to determine the residual signal of the current frame according to: wherein m is a frame index, wherein k is a frequency index, wherein Rm(k) indicates a k-th sample of the residual signal of the current frame in the spectral domain or in the transform domain, wherein Xm(k) indicates a k-th sample of the plurality of spectral coefficients of the current frame in the spectral domain or in the transform domain, wherein indicates a k-th sample of the spectral prediction of the three or more of the plurality of spectral coefficients of the current frame in the spectral domain or in the transform domain, and wherein g is a gain factor. According to an embodiment, the encoder 100 may, e.g., be configured to tabularize and to reduce the computational cost. In the following, a decoder 200 according to particular embodiments is provided. Moreover, a decoder 200 for reconstructing a current frame of an audio signal according to an embodiment is provided. One or more previous frames of the audio signal precede the current frame, wherein each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal, wherein each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain. The decoder 200 is to receive an encoding of the current frame. The decoder 200 is to determine an estimation of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames. The two harmonic parameters for each of the one or more harmonic components of the most previous frame depend on a first group of three or more of the plurality of spectral coefficients of the one or more previous frames of the audio signal. Moreover, the decoder 200 is to reconstruct the current frame depending on the encoding of the current frame and depending on the estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame. The decoder 200 is to receive an encoding of the current frame. One or more previous frames of the audio signal precede the current frame, wherein each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal, wherein each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain. Moreover, the decoder 200 is to determine an estimation of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames. The two harmonic parameters for each of the one or more harmonic components of the most previous frame depend on a first group of three or more of the plurality of spectral coefficients of the one or more previous frames of the audio signal. Furthermore, the decoder 200 is to reconstruct the current frame depending on the encoding of the current frame and depending on the estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame. According to an embodiment, the two harmonic parameters for each of the one or more harmonic components of the most previous frame do not depend on a second group of one or more further spectral coefficients of the plurality of spectral coefficients of the one of more previous frames. In an embodiment, the decoder 200 may, e.g., be to determine an estimation of the two harmonic parameters for each of the one or more harmonic components of the current frame depending on the estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame and depending on the fundamental frequency of the one or more harmonic components of the current frame and the one or more previous frames. According to an embodiment, the decoder 200 may, e.g., be configured to receive the encoding of the current frame comprising a gain factor and a residual signal. The decoder 200 may, e.g., be configured to reconstruct the current frame depending on the gain factor, depending on the residual signal and depending on a fundamental frequency of the one or more harmonic components of the current frame and the one or more previous frames. The fundamental frequency may, e.g., be assumed changing linearly over the current frame and the one or more previous frames. According to an embodiment, the two harmonic parameters for each of the one or more harmonic components are a first parameter for a cosinus sub-component and a second parameter for a sinus sub-component for each of the one or more harmonic components. In an embodiment, the decoder 200 may, e.g., be configured to estimate the two harmonic parameters for each of the one or more harmonic components of the most previous frame by solving a linear equation system comprising at least three equations, wherein each of the at least three equations depends on a spectral coefficient of the first group of the three or more of the plurality of spectral coefficients of the one or more previous frames. According to an embodiment, the decoder 200 may, e.g., be configured to solve the linear equation system using a least mean squares algorithm. According to an embodiment, the linear equations using only the most previous frame are defined by wherein m – 1 is a frame index of the most previous frame, wherein Г is a set of subband indices of length Lm–1, in which the one or more harmonic components lie in the most previous frame, wherein p is a first vector comprising the two harmonic parameters for each of the one or more harmonic components of the most previous frame: wherein ahis a parameter for a cosinus sub-component for an h-th harmonic component of the most previous frame, wherein bhis a parameter for a sinus sub-component for the h-th harmonic component of the most previous frame, wherein H indicates a number of the harmonic components of the most previous frame, wherein Um–1comprises a number of third matrices or third vectors: wherein the third matrix or third vector for an h-th harmonic component of the most previous frame is defined by: wherein Ƴ2his defined by: is the instantaneous angular frequency at time stamp of the h-th harmonic component, wherein is the fundamental frequency of the one or more harmonic components at time stamp , wherein fs is a sampling frequency, wherein Гm–1,his a set of subband indices, in which the h-th harmonic component lie in the most previous frame, wherein is the modulation frequency in subbands defined by Гm–1,h, wherein wherein wherein Bm–1,hrepresents the changing ratio of the angular frequency of the h-th harmonic component in the most previous frame, and for the special case of constant f 0, i.e., B = 0, leads to and equation (25) descends to the identical form of the Modified Discrete Cosine Transform (MDCT) domain representation in [1] for constant fundamental frequencies, wherein f (n) is the (symmetric) window function, wherein N depends on a length of a transform block for transforming the time-domain audio signal into the frequency domain or into the spectral domain. According to an embodiment, linear equations using more previous frames may be added to construct a solvable linear equation system, where the number of previous frames needed is denoted by S. According to an embodiment, when the fundamental frequency of the one or more harmonic components does not change linearly in the current frame and the one or more previous frames, an approximation of a linearly changing fundamental frequency of the one or more harmonic components can be applied in the current frame and the one or more previous frames, where the changing ratio of the angular frequency of the one or more harmonic components may be assumed to be the same in the current frame and the one or more previous frames, wherein the changing ratio of the angular frequency of the h-th harmonic component in the current frame and the one or more previous frames can be obtained by: with s = 1, …, S, wherein f 0mis the fundamental frequency of the one or more harmonic components at the center of the m-th frame. Furthermore, an approximation of the fundamental frequency of the one or more harmonic components in the one or more previous frames can be obtained from a linear extrapolation using the approximated changing ratio of the angular frequency and the fundamental frequencies of the one or more harmonic components in the current frame and the most previous frame: with s = 1, …, S. Alternatively, according to an embodiment, when the fundamental frequency of the one or more harmonic components does not change linearly in the current frame and the one or more previous frames, the fundamental frequency of the one or more harmonic components could be assumed piece-wisely linearly changing in two neighboring frames of the current frame and the one or more previous frames, wherein In an embodiment, the linear equation system is solvable according to: wherein is a first vector comprising an estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame, wherein is a second vector comprising the first group of the three or more of the plurality of spectral coefficients of the one or more previous frames, wherein is a Moore-Penrose inverse matrix of U, wherein U comprises a number of third matrices or third vectors of the one or more previous frames. In an embodiment, the decoder 200 may, e.g., be configured to receive a fundamental frequency of harmonic components, a window function, the gain factor and the residual signal. The decoder 200 may, e.g., be configured to reconstruct the current frame depending on a fundamental frequency of the one or more harmonic components of the most previous frame, depending on the order of the harmonic components, depending on the window function, depending on the gain factor and depending on the residual signal. Only the fundamental frequency, the order of harmonic components, the window function, the gain factor and the residual need to be transmitted. The decoder 200 may, e.g., calculate U based on this received information, and then conduct the harmonic parameters estimation and current frame prediction. The decoder 200 may, e.g., then reconstruct the current frame by adding the transmitted residual spectra to the predicted spectra, scaled by the transmitted gain factor. According to an embodiment, the decoder 200 may, e.g., be configured to receive the number of the one or more harmonic components of the most previous frame and a fundamental frequency of the one or more harmonic components of the most previous frame. The decoder 200 may, e.g., be configured to decode the encoding of the current frame depending on the number of the one or more harmonic components of the most previous frame and depending on the fundamental frequency of the one or more harmonic components of the current frame and the one or more previous frames. According to an embodiment, the decoder 200 is to decode the encoding of the current frame depending on one or more groups of harmonic components, wherein the decoder 200 is to apply a prediction of the audio signal on the one or more groups of harmonic components. According to an embodiment the decoder 200 may, e.g., be configured to determine the two harmonic parameters for each of the one or more harmonic components of the current frame depending on the two harmonic parameters for each of said one of the one or more harmonic components of the most previous frame. According to an embodiment, the decoder 200 may, e.g., be configured to predict the harmonic parameters of the one or more harmonic components in the current frame, denoted as: wherein chis a parameter for a cosinus sub-component for the h-th harmonic component of said one or more harmonic components of the current frame, wherein dhis a parameter for a sinus sub-component for the h-th harmonic component of said one or more harmonic components of the current frame: wherein According to an embodiment, the decoder 200 may, e.g., be configured to receive a residual signal, wherein the residual signal depends on the plurality of spectral coefficients of the current frame in the frequency domain or in the transform domain, and wherein the residual signal depends on the estimation of the two harmonic parameters for each of the one or more harmonic components of the current frame. In an embodiment, the decoder 200 may, e.g., be configured to determine a spectral prediction of one or more of the plurality of spectral coefficients of the current frame depending on the estimation of the two harmonic parameters for each of the one or more harmonic components of the current frame: wherein is a third vector comprising the spectral prediction of the three or more of the plurality of spectral coefficients of the current frame, wherein is a set of subband indices of length Lm, in which the one or more harmonic components lie in the current frame, wherein may be different from Г, when the fundamental frequency of the one or more harmonic components changes between the most previous frame and the current frame. wherein Umcomprises a number of third matrices or third vectors: wherein the third matrix or third vector for an h-th harmonic component of the current frame is defined by: wherein Ƴ2his defined by: is the instantaneous angular frequency at time stamp of the h-th harmonic component, wherein is the fundamental frequency of the one or more harmonic components at time stamp , wherein fs is a sampling frequency, wherein Гm,his a set of subband indices, in which the h-th harmonic component lie in the most previous frame, wherein is the modulation frequency in subbands defined by Гm,h, wherein wherein wherein Bm,hrepresents the changing ratio of the angular frequency of the h-th harmonic component in the current frame, and for the special case of constant f 0, i.e., B = 0, leads to and equation (37) descends to the identical form of the Modified Discrete Cosine Transform (MDCT) domain representation in [1] for constant fundamental frequencies, wherein f (n) is the (symmetric) window function, wherein N depends on a length of a transform block for transforming the time-domain audio signal into the frequency domain or into the spectral domain. According to an embodiment, the decoder 200 may, e.g., be configured to determine the current frame of the audio signal depending on the spectral prediction of the three or more of the plurality of spectral coefficients of the current frame and depending on the residual signal and depending on a gain factor: wherein m is a frame index, wherein k is a frequency index, wherein Rm(k) indicates a k-th sample of the residual signal of the current frame in the spectral domain or in the transform domain, wherein Xm(k) indicates a k-th sample of the plurality of spectral coefficients of the current frame in the spectral domain or in the transform domain, wherein indicates a k-th sample of the spectral prediction of the three or more of the plurality of spectral coefficients of the current frame in the spectral domain or in the transform domain, and wherein g is a gain factor. According to an embodiment, the decoder 200 may, e.g., be configured to decode the audio signal depending on a refined fundamental frequency and depending on an adapted gain factor, which have been determined on a frame basis. According to an embodiment, the encoder 100 (or, e.g., the decoder 200) may, e.g., be configured to tabularize and to reduce the computational cost. In the following, further particular embodiments are provided. A signal model and its frequency domain representation is provided, in which, it is assumed that the fundamental frequency of a tonal signal changes linearly with time, where the linearity can be assumed to be local, and thus be updated across frames, which will be described later with respect to harmonics estimation and prediction below. In the continuous time domain, the phase accumulation of a sinusoid with a linearly changing frequency from time toto time t can be written as where Btrepresents the angular frequency changing ratio. Assume the harmonic components of a tonal signal in the digital time domain with sampling rate fscan be written as below: where n is the sample index in digital time domain, h is the harmonic index, and H is the number of harmonics. According to Equation (45), the phase accumulation of the h-th harmonic from sample n0to sample n can be formulated as: where the angular frequency wh(n) is defined as: where f 0(n) denotes the fundamental frequency at time sampIe n, and Bhdenotes the changing ratio of the normalized angular frequency of the h-th harmonic. Based on above equations, for a block length of 2N of the tonal signal, we deliberately divide the initial phase term of each harmonic component into two parts for the convenience of the later derivation with N as the MDCT frame length and ϕhthe phase reminder. The harmonic components can be represented as: with 0 ≤ n ≤ 2N – 1 and Ch= wh(0) denoting the normalized angular frequency of the h- th harmonic in the first sample of the block. Assuming the frequency information is known, the unknown harmonic parameters are the amplitude Ahand the phase term ϕh,. Equation (49) is rewritten as follows: with unknown harmonic parameters: Based on Equation (50), transforming the block of 2N samples into the MDCT domain results in: with 0 ≤ k ≤ N – 1 and f(n) being the window function. With some derivations, Equation (52) can be rewritten as: where both FCh() and FSh() can be decomposed as the product of a linear phase term and a Fourier transform: For the special case of constant f 0, i.e., B = 0, leads to FS(w) = 0, and Equation (53) descends to the identical form of the MDCT domain representation in [1] for constant fundamental frequencies. In the following, harmonics estimation and prediction according to embodiments are described. In particular, the extended Least Mean Square (LMS) concepts and the prediction concepts according to embodiments are provided. Similar to the FDJHP algorithm [1], to predict an MDCT frame with the proposed EFDJHP algorithm of embodiments, the harmonic parameters in its most previous frame are to be estimated by solving a different linear equation system. To construct a solvable linear equation system, multiple previous frames may be needed. The number of previous frames needed, denoted as S, is dependent on the frequency resolution of the MDCT bins and how dense the harmonic components span on them. Despite the assumption of a linearly changing fundamental frequency in neighboring frames, in the tonal part of real audio signals, the fundamental frequency usually varies with time. An approximation of a linearly changing fundamental frequency can be applied by assuming a constant changing ratio of the angular frequency in the current and S most previous frames, where: = 1, … , S, wherein f 0mis the fundamental frequency at the center of the m-th frame. Furthermore, the fundamental frequency in the S most previous frames can be approximated from a linear extrapolation using Alternatively, the fundamental frequency of the harmonic components could be assumed piece-wisely linearly changing in two neighboring frames of the current frame and the S most frames, wherein In the following, harmonics estimation according to an embodiment is described. Because of the changing frequency, the harmonic energy can be smeared along a few MDCT bins. We denote the joint set of MDCT bins where the harmonics span in the (m-1)-th frame by , with size L. Based on Equation (53), the linear equations using only the most previous frame can be formulated as: Um-1can be calculated when f(n), N and the fundamental frequencies in the (m-1)-th and m-th frames are known. Based on Equation (56), a solvable linear equation system can be built using S most previous frames. The harmonic parameters p can then be estimated by least mean square solving the linear equation system. In the following, prediction according to an embodiment is described. The current MDCT frame is predicted as: where is a joint set of MDCT bins where the harmonics span in the m-th frame and may be different from F if the pitch of the harmonics changes between the (m-1)-th and m-th frame, and denotes a vector of the estimated harmonic parameters in the m-th frame: which can be derived as follows: where Due to nonstationarities in the input signal, the amplitudes of the harmonics may sIightly vary between successive frames. To accommodate with the amplitude variation between frames, a gain factor is introduced to represent these amplitude changes and transmitted as part of the side information to the decoder 200. In the encoding stage, iterations on the gain factor can be done frame wise to find the optimal gain factor. For the bins where no prediction is done, the prediction value is set to zero. The residual spectrum then is: In the following, experiments to evaluate EFDJHP’s improvement over FDJHP and results of the evaluation are provided. For the evaluation of the performance enhancement of EFDJHP in comparison to FDJHP [1], in an experimental set, the transform audio encoder 100 environment implemented in Python, with MDCT as the Time-Frequency transform, and the psychoacoustic model for quantization stepsize calculation [9] have been employed. The sampling rate has been set to 16 KHz. Synthetic sweep signals have first been tested with linearly changing fundamental frequency, and speech and music items with pitch variations have been used. As a reference, the LTP methods in [2] and [4], denoted as AAC LTP (AACLTP) and Frequency Domain Separate Harmonics Prediction (FDSHP) have also been evaluated, respectively. The prediction efficiency has been compared by measuring the Perceptual Entropy (PE) of the prediction residual. PE is an approximation of the signal entropy with consideration of the perceptual model
[0010] . It has been important to have a precise frequency estimate for LTP: It has been noticed that the attacks and decays are not well captured by the YIN estimator
[0011] , which was used previously in
[0011] . Praat is a time domain autocorrelation based pitch estimation algorithm
[0012] and was rated the most accurate pitch estimator an speech signals in
[0013] . The Parselmouth implementation of Praat has been used to obtain the pitch estimates. Similar to FDJHP, EFDJHP can also be performed on selected subbands with limited bandwidth. In the encoder 100 the decision of prediction on subbands can be determined using prediction efficiency measures such as the PE, and signaled in each band to the decoder 200. A comparison of prediction efficiency on sweep signals has been conducted. To test the robustness of EFDJHP against pitch variation, in comparison to FDJHP, synthetic sweep signals with linearly increasing fundamental frequencies have been generated, with the sweep rate varying from 0 Hz / ms to 1.5 Hz / ms. The pitch starts at 100 Hz and goes up to 700 Hz. Bitrate savings of a transform coder using EFDJHP and FDJHP are shown in Fig. 2. In particular, Fig. 2 illustrates an improvement of coding gain using EFDJHP on sweep signals with linearly changing fundamental frequencies, wherein the MDCT frame length is 256. The number of higher harmonics are limited to not exceed the critical bandwidth. To exclude an influence of imprecise pitch information and quantization noise on the prediction efficiency comparison, ground-truth pitches are fed to the LTP unit and no quantization is done on the spectra. For a sweep ratio which equals 0, FDJHP and EFDJHP are equivalent and do not show any performance differences. It can be perceived that FDJHP’s prediction efficiency drops as the sweep rate increases, while EFDJHP’s prediction efficiency remains high despite high sweep rates. The implementation of FDSHP and FDJHP has been improved, compared to the implementation in [1]. In the implementation of EFDJHP, the PE of the residual spectrum in each frame is calculated, with the changing ratio of the fundamental frequency in the neighboring frames being approximated according to Eq.31, and with the changing ratio of the fundamental frequency set to 0. The residual spectrum with the lower PE is taken to ensure that EFDJHP would always perform better or no worse than FDJHP. For conducting a comparison of coding efficiency on speech and music signals a few speech and music items with high frequency variations for comparing coding efficiency on realistic audio signals have been selected. Praat has been used to estimate the fundamental frequencies, and quantization on residual spectra has been enabled, using quantization stepsizes calculated on the input signal using a psychoacoustic model. Iteration around the estimated pitch has been done frame-wise based on the PE. In particular, Praat has been used to estimated the fundamental frequencies, quantization on residual spectra is enabled, using quantization stepsizes calculated on the input signal using a psychoacoustic model [9]. Different MDCT frame lengths (64, 128, and 256) have been used. FDJHP, FDSHP and AACLTP have been used as reference methods. The implementation of FDSHP and FDJHP has been improved. In the implementation of EFDJHP, the PE of the residual spectrum in each frame has been calculated, with the changing ratio of the fundamental frequency in the neighboring frames being approximated according to Equation (11) and with the changing ratio of the fundamental frequency set to 0. The residual spectrum with the lower PE is taken to ensure that EFDJHP would always perform better or no worse than FDJHP. For all working mixes, no prediction has been done in frames, for which no valid pitch estimate has been obtained from Praat, or if the PE of the residual spectrum has been higher than that of the original signal spectrum. The average PE of the quantized residual spectra and of the quantized original signal spectra has then been estimated. Based on the estimated PEs, the Bitrate Saving (BS) by the prediction have been calculated in kbps, without taking the bitrate consumption of side information into account. The coding efficiency of all four LTP methods with and without F0 optimization have all been compared together. For the condition with F0 optimization, a finer pitch search around the Praat estimate and an optimal gain factor search have been done jointly in each frame by minimizing the PE of the quantized residual. For the condition without F0 optimization, an optimal gain factor search has been done in each frame by minimizing the PE of the quantized residual. Fig.3(a), Fig.3(b) and Fig.3(c) illustrate a bitrate saving comparison between AACLTP. FDSHP, FDJHP and EFDJHP, wherein the prediction bandwidth has been selected to be 4 kHz. The results in Fig.3(a)–3(c) show that FDSHP’ performance drops when the frame length is small, where significant overlap of harmonics may exist on the MDCT bins, while AACLTP performs better when the frame length is small. These observations are consistent with that in [1]. In the scenarios with and without F0 optimization, FDJHP always outperforms AACLTP and FDSHP, while EFDJHP performs always the best. For larger frame lengths of 256 (a window length 32 ms), where more pitch variation exists within an MDCT frame, EFDJHP without F0 optimization already outperforms all other methods with F0 optimization on most of the test items, suggesting its superiority in accommodating pitch variations. Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, one or more of the most important method steps may be executed by such an apparatus. Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software or at least partially in hardware or at least partially in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable. Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed. Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine readable carrier. Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier. In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer. A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitory. A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet. A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein. A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein. A further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver. In some embodiments, a programmable logic device (for example a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus. The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer. The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer. The above described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the arrangements and the details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the impending patent claims and not by the specific details presented by way of description and explanation of the embodiments herein. References: [1] Ning Guo and Bernd Edler, “Frequency domain long-term prediction for low delay general audio coding,” IEEE Signal Processing letters, vol. 28, pp. 1185–1189, 2021. [2] J. Ojanperä, M . Väänänen and L. Yin, “Long term predictor for transform domain perceptual audio coding,“ in Audio Engineering Society Convention 107, Sep. 1999. [3] L. Villemoes, J. Klejsa, and P. Hedelin, “Speech coding with transform domain prediction,” in 2017 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2017, pp.324–328. [4] C. Helmrich, Efficient perceptual audio coding using cosine and sine modulated tapped transforms, doctoral thesis, Friedrich-Alexander-Universität Erlangen- Nürnberg (FAU), 2017, Chapter 3.3: Frequency-Domain Prediction with Very Low Complexity. [5] J. Song, C. Lee, H. Oh, and H. Kang, “Harmonic enhancement in low bitrate audio coding using an efficient long-term predictor,“ EURASIP Journal on Advances in Signal Processing, vol.2010, Jan 2010. [6] Goran Markovic, Transform-based Coding Methods for Speech and other Audio Signals, doctoralthesis, Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU), 2022. [7] Bernd Edler, Sascha Disch, Stefan Bayer, Guillaume Fuchs, and Ralf Geiger, “A time-warped MDCT approach to speech transform coding,” Journal of the Audio Engineering Society, May 2009. [8] Tejaswi Nanjundaswamy and Kenneth Rose, “On accommodating pitch variation in long term prediction of speech and vocals in audio coding,” Journal of the Audio Engineering Society, October 2012. [9] A. Taghipour, Psychoacoustics of detection of tonality and asymmetry of masking: implementation of tonality estimation methods in a psychoacoustic model for perceptual audio coding, doctoral thesis, Friedrich-Alexander-Universität Erlangen- Nürnberg (FAU), 2016, Chapter 4: The Psychoacoustic model.
[0010] J. D. Johnston, “Estimation of perceptual entropy using noise masking criteria,” in ICASSP-88,, International Conference on Acoustics, Speech, and Signal Processing, April 1988, pp.2524–2527 vol.5.
[0011] A. de Cheveigné and H. Kawahara, “Yin, a fundamental frequency estimator for speech and music,” The Journal of the Acoustical Society of America, vol.111, pp. 1917–30, 052002.
[0012] Paul Boersma, “Accurate short-term analysis of the fundamental frequency and the harmonics-to-noise ratio of a sampled sound,“ in Proc. Institute of Phonetic Sciences, 1993.
[0013] Sofia Strömbergsson, “Today’s most frequently used f0 estimation methods, and their accuracy in estimating male and female pitch in clean speech,“ 2016.
[0014] Ning Guo and Bernd Edler, “Encoder, decoder, encoding method and decoding method for frequency domain long-term prediction of tonal signals for audio coding," European Patent Application EP 4066242 A1, Oct.2022.
Claims
Claims 1. An encoder (100) for encoding a current frame of an audio signal depending on one or more previous frames of the audio signal, wherein the one or more previous frames precede the current frame, wherein each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal, wherein each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain, wherein the encoder (100) is configured to determine a change of a pitch or an estimation of a change of the pitch between a most previous frame of the one or more previous frames and a current frame, wherein the encoder (100) is configured to generate an encoding of the current frame depending on an estimation of two harmonic parameters for each of one or more harmonic components of the most previous frame, and depending on the change of the pitch or the estimation of the change of the pitch.
2. An encoder (100) according to claim 1, wherein the encoder (100) is configured to determine information on two harmonic parameters of each of one or more harmonic components of the current frame.
3. An encoder (100) according to claim 2, wherein the information on the two harmonic parameters of each of the one or more harmonic components of the current frame depends on the one or more harmonic components of the most previous frame.
4. An encoder (100) according to claim 2 or 3, wherein the encoder (100) is configured to generate the information on the two harmonic parameters of each of the one or more harmonic components of the current frame, such that the two harmonic parameters of each of the one or more harmonic components of the current frame depend on the change of the pitch or the estimation of the change of the pitch.
5. An encoder (100) according to one of the preceding claims, wherein the encoder (100) is configured to operate in a backward adaptive fashion such that reconstructed spectral coefficients are employed in a harmonic estimation step.
6. An encoder (100) according to one of the preceding claims, wherein the encoder (100) is configured to generate the encoding of the current frame such that the encoding of the current frame comprises information on the pitch or the estimation of the change of the pitch.
7. An encoder (100) according to one of the preceding claims, wherein the encoder (100) is configured to generate the encoding of the current frame depending on an assumption that the pitch changes linearly between neighboring frames.
8. An encoder (100) according to one of the preceding claims, wherein the encoder (100) is configured to generate the encoding of the current frame by determining and encoding a gain factor for the current frame.
9. An encoder (100) according to claim 8, wherein the encoder (100) is configured to encode the current frame by determining and encoding a residual signal for the current frame which depends on the gain factor.
10. An encoder (100) according to claim 9, wherein the encoder (100) is configured to determine a refined fundamental frequency and an optimal gain factor on a frame basis by minimizing a perceptual entropy of a quantized version of the residual signal, wherein the encoder (100) is configured to transmit the refined fundamental frequency and the optimal gain factor to a decoder.
11. An encoder (100) according to claim 9 or 10, wherein the encoder (100) is configured to generate the encoding of the current frame by determining the gain factor and the residual signal depending on a fundamental frequency of the one or more harmonic components of the current frame and depending on a fundamental frequency of each of the one or more previous frames.
12. An encoder (100) according to claim 11, wherein the encoder (100) is configured to generate the encoding of the current frame by determining the gain factor and the residual signal depending on the estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame.
13. An encoder (100) according to one of the preceding claims, wherein the encoder (100) is to determine the estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame.
14. An encoder (100) according to claim 13, wherein the encoder (100) is to determine the estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame using a first group of three or more of the plurality of spectral coefficients of the one or more previous frames of the audio signal.
15. An encoder (100) according to one of the preceding claims, wherein the encoder (100) is configured to determine an estimation of two harmonic parameters for each of the one or more harmonic components of the current frame depending on an estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame and depending on a fundamental frequency of the one or more harmonic components of the current frame and depending on a fundamental frequency of each of the one or more previous frames.
16. An encoder (100) according to claim 15, wherein the encoder (100) is configured to determine the estimation of two harmonic parameters for each of the one or more harmonic components of the current frame depending on an assumption that the fundamental frequency changes linearly over the current frame and the one or more previous frames.
17. An encoder (100) according to one of the preceding claims, wherein the two harmonic parameters for each of the one or more harmonic components are a first parameter for a cosinus sub-component and a second parameter for a sinus sub-component for each of the one or more harmonic components.
18. An encoder (100) according to one of the preceding claims, wherein the encoder (100) is configured to estimate the two harmonic parameters for each of the one or more harmonic components of the most previous frame by solving a linear equation system comprising at least three equations.
19. An encoder (100) according to claim 18, wherein each of the at least three equations depends on a spectral coefficient of the first group of the three or more of the plurality of spectral coefficients of the one or more previous frames.
20. An encoder (100) according to claim 18 or 19, wherein the encoder (100) is configured to solve the linear equation system using a least mean squares algorithm.
21. An encoder (100) according to one of claims 18 to 20, wherein the at least three linear equations are defined bywherein m – 1 is a frame index of the most previous frame, wherein Г is a set of subband indices of length Lm–1, in which the one or more harmonic components lie in the most previous frame, wherein p is a first vector comprising the two harmonic parameters for each of the one or more harmonic components of the most previous frame, wherein Um–1comprises a number of third matrices or third vectors.
22. An encoder (100) according to claim 21, whereinwherein ahis a parameter for a cosinus sub-component for an h-th harmonic component of the most previous frame, wherein bhis a parameter for a sinus sub-component for the h-th harmonic component of the most previous frame, wherein H indicates a number of the harmonic components of the most previous frame.
23. An encoder (100) according to claim 21 or 22, wherein.
24. An encoder (100) according to claim 23, wherein a third matrix or third vector for an h-th harmonic component of the most previous frame is defined by:wherein Ƴ2his defined by:is an instantaneous angular frequency at time stamp of the h-th harmonic component, whereinis the fundamental frequency of the one or more harmonic components at time stamp , wherein fs is a sampling frequency, wherein Гm–1,his a set of subband indices, in which the h-th harmonic component lie in the most previous frame, whereinis a modulation frequency in subbands defined by Гm–1,h.
25. An encoder (100) according to claim 24, whereinwhereinwhereinrepresents the changing ratio of the angular frequency of the h-th harmonic component in the most previous frame, wherein f (n) is a window function, wherein N depends on a length of a transform block for transforming the time- domain audio signal into the frequency domain or into the spectral domain.
26. An encoder (100) according to one of claims 18 to 25, wherein the encoder (100) is configured to add one or more further linear equations using one or more further previous frames to the linear equation system to construct a solvable linear equation system.
27. An encoder (100) according to one of the preceding claims, wherein, to generate the encoding of the current frame, when the fundamental frequency of the one or more harmonic components does not change linearly in the current frame and the one or more previous frames, the encoder (100) is configured to apply an approximation of a linearly changing fundamental frequency of the one or more harmonic components in the current frame and in the one or more previous frames.
28. An encoder (100) according to one of the preceding claims, wherein, to generate the encoding of the current frame, the encoder (100) is configured to obtain the changing ratio of the angular frequency of the h-th harmonic component in the current frame and the one or more previous frames by:with s = 1, …, S, where S denoted the number of previous frames needed, wherein f0mis the fundamental frequency of the one or more harmonic components at the center of the m-th frame.
29. An encoder (100) according to one of the preceding claims, wherein, to generate the encoding of the current frame, the encoder (100) is configured to obtain an approximation of a fundamental frequency of the one or more harmonic components in the one or more previous frames by conducting a linear extrapolation using an approximated changing ratio of the angular frequency and the fundamental frequencies of the one or more harmonic components in the current frame and the most previous frame.
30. An encoder (100) according to claim 29, wherein the approximated changing ratio of the angular frequency and the fundamental frequencies of the one or more harmonic components in the current frame and the most previous frame is determined as:with s = 1, …, S, where S denoted the number of previous frames needed, wherein f0mis the fundamental frequency of the one or more harmonic components at the center of the m-th frame.
31. An encoder (100) according to one of the preceding claims,wherein, to generate the encoding of the current frame, when the fundamental frequency of the one or more harmonic components does not change linearly in the current frame and the one or more previous frames, the encoder (100) is configured to assume that the fundamental frequency of the one or more harmonic components is changing piece-wisely linearly in two neighboring frames of the current frame and the one or more previous frames.
32. An encoder (100) according to claim 31, whereinwith s = 1, …, S, where S denoted the number of previous frames needed, wherein f0mis the fundamental frequency of the one or more harmonic components at the center of the m-th frame.
33. An encoder (100) according to one of the preceding claims, wherein, to generate the encoding of the current frame the encoder (100) is configured to solve a linear equation system: wherein is a first vector comprising an estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame, wherein is a second vector comprising the first group of the three or more of the plurality of spectral coefficients of the one or more previous frames, wherein is a Moore-Penrose inverse matrix of U, wherein U comprises a number of third matrices or third vectors of the one or more previous frames.
34. An encoder (100) according to one of the preceding claims, wherein the encoder (100) is configured to encode a fundamental frequency of harmonic components, a window function, a gain factor and a residual signal.
35. An encoder (100) according to one of the preceding claims, wherein, to generate the encoding of the current frame, the encoder (100) is configured to determine the number of the one or more harmonic components of the most previous frame and a fundamental frequency of the one or more harmonic components of the most previous frame before estimating the two harmonic parameters for each of the one or more harmonic components of the most previous frame using a first group of three or more of the plurality of spectral coefficients for each of the one or more previous frames of the audio signal.
36. An encoder (100) according to one of the preceding claims, wherein, to generate the encoding of the current frame, the encoder (100) is configured to predict the harmonic parameters of the one or more harmonic components in the current frame aswherein chis a parameter for a cosinus sub-component for the h-th harmonic component of said one or more harmonic components of the current frame, wherein dhis a parameter for a sinus sub-component for the h-th harmonic component of said one or more harmonic components of the current frame.
37. An encoder (100) according to claim 36, whereinwherein38. An encoder (100) according to one of the preceding claims, wherein, to generate the encoding of the current frame, the encoder (100) is configured to determine a spectral prediction of one or more of the plurality of spectral coefficients of the current frame depending on the estimation of the two harmonic parameters for each of the one or more harmonic components of the current frame.
39. An encoder (100) according to one of the preceding claims, wherein the encoder (100) is configured to determine the spectral predictionwherein is a third vector comprising the spectral prediction of the three or more of the plurality of spectral coefficients of the current frame, wherein is a set of subband indices of length Lm, in which the one or more harmonic components lie in the current frame, wherein may be different from Г, when the fundamental frequency of the one or more harmonic components changes between the most previous frame and the current frame, wherein Umcomprises a number of third matrices or third vectors:.
40. An encoder (100) according to claim 39, wherein the third matrix or third vector for an h-th harmonic component of the current frame is defined by:wherein Ƴ2his defined by:is an instantaneous angular frequency at time stamp of the h-th harmonic component, whereinis a fundamental frequency of the one or more harmonic components at time stamp , wherein fs is a sampling frequency, wherein Гm,his a set of subband indices, in which the h-th harmonic component lie in the most previous frame, whereinis the modulation frequency in subbands defined by Гm,h.
41. An encoder (100) according to claim 40, whereinwhereinwherein Bm,hrepresents the changing ratio of the angular frequency of the h-th harmonic component in the current frame, wherein f (n) is a window function, wherein N depends on a length of a transform block for transforming the time- domain audio signal into the frequency domain or into the spectral domain.
42. An encoder (100) according to one of the preceding claims, wherein, to generate the encoding of the current frame, the encoder (100) is configured to determine a residual signal and a gain factor depending on the plurality of spectral coefficients of the current frame in a frequency domain or in a transform domain and depending on a spectral prediction of the three or more of the plurality of spectral coefficients of the current frame, wherein the encoder (100) is configured to generate the encoding of the current frame such that the encoding of the current frame comprises the residual signal and the gain factor.
43. An encoder (100) according to one the preceding claims, wherein the encoder (100) is configured to transmit only the fundamental frequency, the order of harmonic components, the window function, the gain factor and the residual to a decoder (200).
44. An encoder (100) according to one of the preceding claims,wherein, to generate the encoding of the current frame, the encoder (100) is configured to determine a residual signal of the current frame according to:wherein m is a frame index, wherein k is a frequency index, wherein Rm(k) indicates a k-th sample of the residual signal of the current frame in a spectral domain or in a transform domain, wherein Xm(k) indicates a k-th sample of the plurality of spectral coefficients of the current frame in the spectral domain or in the transform domain, wherein indicates a k-th sample of the spectral prediction of the three or more of the plurality of spectral coefficients of the current frame in the spectral domain or in the transform domain, and wherein g is a gain factor.
45. An encoder (100) according to one of the preceding claims, wherein, to generate the encoding of the current frame, the encoder (100) is configured to tabularize.
46. A decoder (200) for reconstructing a current frame of an audio signal, wherein one or more previous frames of the audio signal precede the current frame, wherein each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal, wherein each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain, wherein the decoder (200) is to receive an encoding of the current frame,wherein the decoder (200) is to determine an estimation of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames, wherein the decoder (200) is to reconstruct the current frame depending on the encoding of the current frame, wherein the decoder (200) is to reconstruct the current frame depending on the estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame, wherein the decoder (200) is to reconstruct the current frame depending on a change of a pitch or an estimation of a change of the pitch between a most previous frame of the one or more previous frames and a current frame.
47. A decoder (200) according to claim 46, wherein the decoder (200) is to reconstruct the current frame by determining information on two harmonic parameters of each of one or more harmonic components of the current frame.
48. A decoder (200) according to claim 47, wherein the information on the two harmonic parameters of each of the one or more harmonic components of the current frame depends on the one or more harmonic components of the most previous frame.
49. A decoder (200) according to claim 47 or 48, wherein the two harmonic parameters of each of the one or more harmonic components of the current frame depend on the change of the pitch or the estimation of the change of the pitch.
50. A decoder (100) according to one of claims 46 to 49, wherein the decoder (200) is configured to operate in a backward adaptive fashion such that reconstructed spectral coefficients are employed in a harmonic estimation step.
51. A decoder (200) according to one of claims 46 to 50, wherein the encoding of the current frame comprises information on the pitch or the estimation of the change of the pitch, wherein the decoder (200) is to reconstruct the current frame depending on the estimation of the change of the pitch.
52. A decoder (200) according to one of claims 46 to 51, wherein the encoding of the current frame depends on an assumption that the pitch changes linearly between neighboring frames.
53. A decoder (200) according to one of claims 46 to 52, wherein the encoding of the current frame comprises a gain factor for the current frame, wherein the decoder (200) is to reconstruct the current frame depending on the gain factor for the current frame.
54. A decoder (200) according to claim 53, wherein the encoding of the current frame comprises a residual signal for the current frame which depends on the gain factor, wherein the decoder (200) is to reconstruct the current frame depending on the residual signal for the current frame.
55. A decoder (200) according to claim 54, wherein the decoder (200) is configured to receive a refined fundamental frequency and an optimal gain factor, wherein the decoder (200) is to reconstruct the current frame depending on the refined fundamental frequency and depending on the optimal gain factor.
56. A decoder (200) according to claim 54 or 55, wherein the gain factor and the residual signal depend on a fundamental frequency of the one or more harmonic components of the current frame and depend on a fundamental frequency of each of the one or more previous frames.
57. A decoder (200) according to claim 56, wherein the gain factor and the residual signal depend on the estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame.
58. A decoder (200) according to one of claims 46 to 57, wherein the two harmonic parameters for each of the one or more harmonic components of the most previous frame depend on a first group of three or more of the plurality of spectral coefficients of the one or more previous frames of the audio signal.
59. A decoder (200) according to one of claims 46 to 58, wherein the two harmonic parameters for each of the one or more harmonic components are a first parameter for a cosinus sub-component and a second parameter for a sinus sub-component for each of the one or more harmonic components.
60. A decoder (200) according to one of claims 46 to 59, wherein the decoder (200) is configured to estimate the two harmonic parameters for each of the one or more harmonic components of the most previous frame by solving a linear equation system comprising at least three equations.
61. A decoder (200) according to claim 60, wherein each of the at least three equations depends on a spectral coefficient of the first group of the three or more of the plurality of spectral coefficients of the one or more previous frames.
62. A decoder (200) according to claim 61, wherein the decoder (200) is configured to solve the linear equation system using a least mean squares algorithm.
63. A decoder (200) according to one of claims 60 to 62, wherein the at least three linear equations are defined bywherein m – 1 is a frame index of the most previous frame, wherein Г is a set of subband indices of length Lm–1, in which the one or more harmonic components lie in the most previous frame, wherein p is a first vector comprising the two harmonic parameters for each of the one or more harmonic components of the most previous frame, wherein Um–1comprises a number of third matrices or third vectors.
64. A decoder (200) according to claim 63, whereinwherein ahis a parameter for a cosinus sub-component for an h-th harmonic component of the most previous frame, wherein bhis a parameter for a sinus sub-component for the h-th harmonic component of the most previous frame, wherein H indicates a number of the harmonic components of the most previous frame.
65. A decoder (200) according to claim 63 or 64, wherein.
66. A decoder (200) according to claim 65, wherein a third matrix or third vector for an h-th harmonic component of the most previous frame is defined by:wherein Ƴ2his defined by:is an instantaneous angular frequency at time stamp of the h-th harmonic component, wherein isfundamental frequency of the one or more harmoniccomponents at time stamp , wherein fs is a sampling frequency, whereinis a set of subband indices, in which the h-th harmonic component lie in the most previous frame.
67. A decoder (200) according to claim 66, whereinis the modulation frequency in subbands defined by Гm–1,h, whereinwhereinwherein Bm–1,hrepresents the changing ratio of the angular frequency of the h-th harmonic component in the most previous frame, wherein f (n) is a window function, wherein N depends on a length of a transform block for transforming the time- domain audio signal into the frequency domain or into the spectral domain.
68. A decoder (200) according to one of claims 63 to 67, wherein the decoder (200) is configured to add one or more further linear equations using one or more further previous frames to the linear equation system to construct a solvable linear equation system.
69. A decoder (200) according to one of claims 46 to 68,wherein, when the fundamental frequency of the one or more harmonic components does not change linearly in the current frame and the one or more previous frames, the decoder (200) is configured to apply an approximation of a linearly changing fundamental frequency of the one or more harmonic components in the current frame and the one or more previous frames.
70. A decoder (200) according to one of claims 46 to 69, wherein the decoder (200) is configured to obtain the changing ratio of the angular frequency of the h-th harmonic component in the current frame and the one or more previous frames by:with s = 1, …, S, where S denoted the number of previous frames needed, whereinis the fundamental frequency of the one or more harmonic components at the center of the m-th frame, wherein the decoder (200) is to reconstruct the current frame depending on the changing ratio.
71. A decoder (200) according to one of claims 46 to 70, wherein the decoder (200) is configured to obtain an approximation of the fundamental frequency of the one or more harmonic components in the one or more previous frames can be obtained from a linear extrapolation using the approximated changing ratio of the angular frequency and the fundamental frequencies of the one or more harmonic components in the current frame and the most previous frame, wherein the decoder (200) is to reconstruct the current frame depending on the approximation of the fundamental frequency.
72. A decoder (200) according to claim 71,wherein the approximated changing ratio of the angular frequency and the fundamental frequencies of the one or more harmonic components in the current frame and the most previous frame is determined as:with s = 1, …, S, where S denoted the number of previous frames needed is denoted by wherein f0mis the fundamental frequency of the one or more harmonic components at the center of the m-th frame.
73. A decoder (200) according to one of claims 46 to 72, wherein, when the fundamental frequency of the one or more harmonic components does not change linearly in the current frame and the one or more previous frames, the decoder (200) is configured to assume that the fundamental frequency of the one or more harmonic components is changing piece-wisely linearly in two neighboring frames of the current frame and the one or more previous frames.
74. A decoder (200) according to claim 73, whereinwith s = 1, …, S, where S denoted the number of previous frames needed, wherein f0mis the fundamental frequency of the one or more harmonic components at the center of the m-th frame.
75. A decoder (200) according to one of claims 46 to 74, wherein the decoder (200) is configured to determine a solution of the linear equation system:wherein is a first vector comprising an estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame, whereinis a second vector comprising the first group of the three or more of the plurality of spectral coefficients of the one or more previous frames, wherein is a Moore-Penrose inverse matrix of U, wherein U comprises a number of third matrices or third vectors of the one or more previous frames, wherein the decoder (200) is to reconstruct the current frame depending on the solution.
76. A decoder (200) according to one of claims 46 to 75, wherein the decoder (200) is configured to receive a fundamental frequency of harmonic components, a window function, a gain factor and a residual signal, and wherein the decoder (200) is to reconstruct the current frame depending on the fundamental frequency of harmonic components, the window function, the gain factor and the residual signal.
77. A decoder (200) according to claim 76, wherein the decoder (200) is configured to reconstruct the current frame depending on a fundamental frequency of the one or more harmonic components of the most previous frame, depending on an order of the harmonic components, depending on the window function, depending on the gain factor and depending on the residual signal.
78. A decoder (200) according to one of claims 46 to 77, wherein the decoder (200) is configured to receive only the fundamental frequency, the order of harmonic components, the window function, the gain factor and the residual from an encoder (100).
79. A decoder (200) according to one of claims 46 to 78, wherein the decoder (200) is configured to calculate U based on this received information, and then conduct the harmonic parameters estimation and current frame prediction, wherein the decoder (200) then reconstructs the current frame by adding transmitted residual spectra to predicted spectra, scaled by the transmitted gain factor.
80. A decoder (200) according to one of claims 46 to 79, wherein the decoder (200) is configured to receive a number of the one or more harmonic components of the most previous frame and a fundamental frequency of the one or more harmonic components of the most previous frame, wherein the decoder (200) is configured to decode the encoding of the current frame depending on the number of the one or more harmonic components of the most previous frame and depending on the fundamental frequency of the one or more harmonic components of the current frame and the one or more previous frames.
81. A decoder (200) according to one of claims 46 to 80, wherein the decoder (200) is to decode the encoding of the current frame depending on one or more groups of harmonic components, wherein the decoder (200) is to apply a prediction of the audio signal on the one or more groups of harmonic components.
82. A decoder (200) according to one of claims 46 to 81, wherein the decoder (200) is configured to determine the two harmonic parameters for each of the one or more harmonic components of the current frame dependingon the two harmonic parameters for each of said one of the one or more harmonic components of the most previous frame.
83. A decoder (200) according to one of claims 46 to 82, wherein the decoder (200) is configured to predict the harmonic parameters of the one or more harmonic components in the current frame aswherein chis a parameter for a cosinus sub-component for the h-th harmonic component of said one or more harmonic components of the current frame, wherein dhis a parameter for a sinus sub-component for the h-th harmonic component of said one or more harmonic components of the current frame.
84. A decoder (200) according to claim 83, whereinwherein85. A decoder (200) according to one of claims 46 to 84, wherein the decoder (200) is configured to receive a residual signal, wherein the residual signal depends on the plurality of spectral coefficients of the current frame in a frequency domain or in a transform domain, and wherein the residual signal depends on the estimation of the two harmonic parameters for each of the one or more harmonic components of the current frame.
86. A decoder (200) according to one of claims 46 to 85,wherein the decoder (200) is configured to determine a spectral prediction of one or more of the plurality of spectral coefficients of the current frame depending on the estimation of the two harmonic parameters for each of the one or more harmonic components of the current frame:wherein is a third vector comprising the spectral prediction of the three or more of the plurality of spectral coefficients of the current frame, wherein is a set of subband indices of length Lm, in which the one or more harmonic components lie in the current frame, wherein may be different from Г, when the fundamental frequency of the one or more harmonic components changes between the most previous frame and the current frame, wherein Umcomprises a number of third matrices or third vectors:
87. A decoder (200) according to claim 86, wherein the third matrix or third vector for an h-th harmonic component of the current frame is defined by:wherein Ƴ2his defined by:is an instantaneous angular frequency at time stampof the h-th harmonic component, whereinis the fundamental frequency of the one or more harmonic components at time stamp , wherein fs is a sampling frequency, whereinis a set of subband indices, in which the h-th harmonic component lie in the most previous frame, whereinis the modulation frequency in subbands defined by Гm,h.
88. A decoder (200) according to claim 87, whereinwhereinwherein Bm,hrepresents the changing ratio of the angular frequency of the h-th harmonic component in the current frame, wherein f (n) is a window function, wherein N depends on a length of a transform block for transforming the time- domain audio signal into the frequency domain or into the spectral domain.
89. A decoder (200) according to one of claims 46 to 88, wherein the decoder (200) is configured to determine the current frame of the audio signal depending on the spectral prediction of the three or more of the plurality of spectral coefficients of the current frame and depending on the residual signal and depending on a gain factor:wherein m is a frame index, wherein k is a frequency index, wherein Rm(k) indicates a k-th sample of the residual signal of the current frame in the spectral domain or in the transform domain, wherein Xm(k) indicates a k-th sample of the plurality of spectral coefficients of the current frame in the spectral domain or in the transform domain, whereinindicates a k-th sample of the spectral prediction of the three or more of the plurality of spectral coefficients of the current frame in the spectral domain or in the transform domain, and wherein g is a gain factor.
90. A decoder (200) according to one of claims 46 to 89, wherein the decoder (200) is configured to decode the audio signal depending on a refined fundamental frequency and depending on an adapted gain factor, which have been determined on a frame basis.
91. A decoder (200) according to one of claims 46 to 90, wherein the decoder (200) is configured to tabularizeand to reduce the computational cost.
92. A system, comprising: an encoder (100) according to one of claims 1 to 45, and a decoder (200) according to one of claims 46 to 91, wherein the decoder (200) is configured to receive the encoding of the current frame from the encoder (100).
93. A method for encoding a current frame of an audio signal depending on one or more previous frames of the audio signal, wherein the one or more previous frames precede the current frame, wherein each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal, wherein each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain, wherein the method comprises: determining a change of a pitch or an estimation of a change of the pitch between a most previous frame of the one or more previous frames and a current frame, and generating an encoding of the current frame depending on an estimation of two harmonic parameters for each of one or more harmonic components of the most previous frame, and depending on the change of the pitch or the estimation of the change of the pitch.
94. A method for reconstructing a current frame of an audio signal, wherein one or more previous frames of the audio signal precede the current frame, wherein each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal, wherein each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain, wherein the method comprises:receiving an encoding of the current frame, determining an estimation of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames, and reconstructing the current frame depending on the encoding of the current frame, wherein reconstructing the current frame is conducted depending on the estimation of the two harmonic parameters for each of the one or more harmonic components of the most previous frame, and wherein reconstructing the current frame is conducted depending on a change of a pitch or an estimation of a change of the pitch between a most previous frame of the one or more previous frames and a current frame.
95. A computer program for implementing the method of claim 93 or 94 when being executed on a computer or signal processor.
Citation Information
Patent Citations
Encoder, decoder, encoding method and decoding method for frequency domain long-term prediction of tonal signals for audio coding
EP4066242A1
Method and apparatus for encoding / decoding media signal
CN101790887B
Encoder, decoder, encoding method and decoding method for frequency domain long-term prediction of tonal signals for audio coding
US20220284908A1