Model-based prediction in critically sampled filterbanks
By employing signal model-based subband predictors directly in the subband domain, the method addresses the inefficiencies of low-bitrate audio coding, reducing aliasing artifacts and complexity in critically sampled filter banks.
Patent Information
- Application Number
- JP2025173314
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2013-09-09
- Filing Date
- 2025-10-15
- Publication Date
- 2026-01-27
AI Technical Summary
Existing audio source coding systems face challenges in implementing subband predictors efficiently at low bitrates while minimizing aliasing artifacts and computational complexity, particularly in critically sampled filter banks.
The method involves using a compact description of subband predictors based on signal models, directly implementing predictors in the subband domain, and employing cross-subband terms to reduce aliasing artifacts, allowing for efficient low-bitrate audio coding.
This approach reduces aliasing artifacts and computational complexity, enabling high-quality audio coding at lower bitrates by utilizing sinusoidal frequencies and periodic signal models for periodic and polyphonic signals, and providing seamless transitions between models.
Smart Images

Figure 2026012753000001_ABST
Abstract
Description
[Technical Field]
[0001] This paper relates to an audio source coding system, and in particular to an audio source coding system that utilizes linear prediction in combination with a filter bank. [Background technology]
[0002] There are two important signal processing tools applied in systems for source coding of audio signals: critically sampled filter banks and linear prediction. Critically sampled filter banks (e.g., modified discrete cosine transform (MDCT)-based transforms) allow direct access to time-frequency representations where perceptual insignificance and signal redundancy can be exploited. Linear prediction allows efficient source modeling of audio signals, especially speech signals. The combination of the two tools, i.e., the use of prediction in subbands of filter banks, has been primarily used for high-bitrate audio coding. For low-bitrate coding, the challenge with prediction in subbands is to keep the cost (i.e., bitrate) of the prediction description low. Another challenge is to control the resulting noise shaping of the prediction error signal obtained by the subband predictor.
[0003] Regarding the problem of encoding subband predictor descriptions in a bit-efficient manner, a possible approach is to estimate the predictor from previously decoded portions of the audio signal, thereby avoiding the cost of the predictor description altogether. If the predictor can be determined from previously decoded portions of the audio signal, the predictor can be determined at the encoder and decoder without the need to transmit the predictor description from the encoder to the decoder. This approach is referred to as a backward-adaptive prediction scheme. However, backward-adaptive prediction schemes typically degrade significantly as the bit rate of the encoded audio signal decreases. An alternative or additional approach to efficient encoding of subband predictors is to identify more natural predictor descriptions, e.g., descriptions that exploit the intrinsic structure of the audio signal to be encoded. For example, low-bit-rate speech coding typically applies a forward-adaptive scheme based on a compact representation of a short-term predictor (exploiting short-term correlations) and a long-term predictor (exploiting long-term correlations due to the underlying pitch of the speech signal).
[0004] Regarding the issue of controlling the noise shape of the prediction error signal, it is observed that although the noise shape of the predictor can be well controlled within a subband, the final output audio signal of the encoder typically exhibits aliasing artifacts (except for audio signals that exhibit a substantially flat spectral noise shape).
[0005] An important example of a subband predictor is the implementation of long-term prediction in a filter bank with overlapping windows. LTPs typically exploit redundancies in periodic and near-periodic audio signals (e.g., speech signals that exhibit inherent pitch) and can be described by a single or a few prediction parameters. A LTP can be defined in continuous time by a delay that reflects the periodicity of the audio signal. When this delay is large compared to the filter bank window length, the LTP can be implemented in the discrete time domain by a shift or fractional delay and then transformed back into a causal predictor in the subband domain. While such LTPs typically do not exhibit aliasing artifacts, they incur a significant computational penalty due to the need for additional filter bank operations for the transformation from the time domain to the subband domain. Furthermore, the approach of determining the delay in the time domain and transforming it into a subband predictor is not applicable when the period of the audio signal to be encoded is comparable to or smaller than the filter bank window size. Summary of the Invention [Problem to be solved by the invention]
[0006] This paper addresses the above-mentioned shortcomings of subband prediction. In particular, this paper describes methods and systems that allow for a bitrate-efficient implementation of subband predictors and / or allow for a reduction in aliasing artifacts caused by subband predictors. In particular, the methods and systems described herein enable the implementation of low-bitrate audio coders using subband prediction that result in reduced levels of aliasing artifacts. [Means for solving the problem]
[0007] This paper describes a method and system for improving the quality of audio source coding using prediction in the subband domain of a critically sampled filter bank. The method and system may utilize a compact description of the subband predictor based on a signal model. Alternatively or additionally, the method and system may utilize an efficient implementation of the predictor directly in the subband domain. Alternatively or additionally, the method and system may utilize the cross-subband predictor term described herein to allow for the reduction of aliasing artifacts.
[0008] As outlined in this paper, a compact description of a subband predictor may include sinusoidal frequencies, periods for periodic signals, slightly inharmonic spectra such as those encountered for rigid string vibrations, and / or multiple pitches for polyphonic signals. For the case of long-term predictors, periodic signal models are shown to provide high-quality causal predictors for a range of lag parameters (or delays), including values shorter and / or longer than the filter bank window size. This means that periodic signal models can be used to implement long-term subband predictors in an efficient manner. A seamless transition from sinusoidal model-based prediction to approximations of arbitrary delays is provided. Direct implementation of the predictor in the subband domain allows explicit access to the perceptual properties of the resulting quantization distortion. Furthermore, implementing the predictor in the subband domain allows access to numerical properties such as the prediction gain and the dependence of the prediction on the parameters. For example, signal model-based analysis may reveal that prediction gain is significant only in a subset of the subbands considered, and the variation of predictor coefficients as a function of the parameters chosen for transmission may aid in the design of efficient encoding algorithms as well as parameter formats. Furthermore, the computational complexity may be significantly reduced compared to predictor implementations that rely on the use of algorithms that operate in both the time and subband domains. In particular, the methods and systems described herein may be used to implement subband prediction directly in the subband domain, without the need to determine and apply a predictor (e.g., long-term delay) in the time domain.
[0009] The use of cross-subband terms in subband predictors allows for significantly improved frequency-domain noise shaping properties compared to intraband predictors (which rely solely on intraband prediction). By doing so, aliasing artifacts can be reduced, thereby enabling the use of subband prediction for relatively low bitrate audio coding systems.
[0010] According to one aspect, a method for estimating a first sample of a first subband of an audio signal is described. The first subband of the audio signal may be determined using an analysis filterbank having a plurality of analysis filters that provide a plurality of subband signals in a plurality of subbands from the audio signal. A time-domain audio signal may be submitted to the analysis filterbank, thereby providing a plurality of subband signals in a plurality of subbands. Each of the plurality of subbands typically covers a different frequency range of the audio signal, thereby providing access to different frequency components of the audio signal. The plurality of subbands may have equal or uniform subband spacing. The first subband corresponds to one of the plurality of subbands provided by the analysis filterbank.
[0011] Analysis filter banks may have various attributes. Synthesis filter banks with multiple synthesis filters may have similar or identical attributes. The attributes described for analysis filter banks and analysis filters are also applicable to the attributes of synthesis filter banks and synthesis filters. Typically, the combination of the analysis filter bank and synthesis filter bank allows perfect reconstruction of the audio signal. The analysis filters of an analysis filter bank may be shift-invariant with respect to each other. Alternatively or additionally, the analysis filters of an analysis filter bank may have a common window function. In particular, the analysis filters of an analysis filter bank may have differently modulated versions of a common window function. In one embodiment, the common window function is modulated using a cosine function, thereby providing a cosine-modulated analysis filter bank. In particular, the analysis filter bank may have (or correspond to) one or more of an MDCT, a QMF, and / or an ELT transform. The common window function may have a finite duration K. The duration of the common window function may be such that successive samples of the subband signals are determined using overlapping segments of the time-domain audio signal. Thus, the analysis filterbank may include overlapped transforms. The analysis filters of the analysis filterbank may form an orthogonal and / or orthonormal basis. As a further attribute, the analysis filterbank may correspond to a critically sampled filterbank. In particular, the number of samples of the plurality of subband signals may correspond to the number of samples of the time-domain audio signal.
[0012] The method may include determining model parameters of a signal model. The signal model may be described using a plurality of model parameters. Thus, the method may include determining the plurality of model parameters of the signal model. The model parameter(s) may be extracted from a received bitstream that includes or represents the model parameters and a prediction error signal. Alternatively, the model parameter(s) may be determined by fitting the signal model to the audio signal (e.g., frame by frame), for example using a mean square error approach.
[0013] The signal model may include one or more sinusoidal model components. In such a case, the model parameters may indicate the one or more frequencies of the one or more sinusoidal model components. By way of example, the model parameters may indicate a fundamental frequency Ω of a multi-sinusoidal signal model, where the multi-sinusoidal signal has sinusoidal model components at frequencies corresponding to multiples qΩ of the fundamental frequency Ω. Thus, the multi-sinusoidal signal model may have a periodic signal component, where the periodic signal component includes multiple sinusoidal components, the multiple sinusoidal components having frequencies that are multiples of the fundamental frequency Ω. As described herein, such periodic signal components may be used to model delays in the time domain (e.g., as used for long-term predictors). The signal model may have one or more model parameters indicating a shift and / or deviation of the signal model from a periodic signal model. The shift and / or deviation may indicate a deviation of the frequencies of the multiple sinusoidal components of the periodic signal model from respective multiples qΩ of the fundamental frequency Ω.
[0014] The signal model may have a plurality of periodic signal components, each of which may be described using one or more model parameters, which may represent a plurality of fundamental frequencies Ω0, Ω1, ..., Ω M-1Alternatively or additionally, the signal model may be described by a predetermined and / or adjustable relaxation parameter (which may be one of the model parameters). The relaxation parameter may be configured to even out or smooth the line spectrum of the periodic signal component. Specific examples of signal models and associated model parameters are described in the Examples section of this document.
[0015] The model parameter(s) may be determined such that an average value of a squared prediction error signal is small (e.g., minimized). The prediction error signal may be determined based on a difference between a first sample and an estimate of the first sample. In particular, an average value of the squared prediction error signal may be determined based on a plurality of consecutive first samples of the first subband signal and a corresponding plurality of estimated first samples. In particular, it is proposed to model the audio signal or at least the first subband signal of the audio signal using a signal model described by one or more model parameters. The model parameters are used to determine the one or more prediction coefficients of a linear predictor that determines a first estimated subband signal. The difference between the first subband signal and the first estimated subband signal provides a prediction error subband signal. The one or more model parameters may be determined such that an average value of the squared prediction error subband signal is small (e.g., minimized).
[0016] The method may further include determining prediction coefficients to be applied to previous samples of a first decoded subband signal derived from the first subband signal. In particular, the previous samples may be determined by adding (a quantized version of) the prediction error signal to corresponding samples of the first subband signal. The first decoded subband signal may be identical to the first subband signal (e.g., in the case of a lossless encoder). The time slot of the previous sample is typically earlier than the time slot of the first sample. In particular, the method may include determining one or more prediction coefficients of a recursive (finite impulse response) prediction filter configured to determine the first sample of the first subband signal from one or more previous samples.
[0017] The one or more prediction coefficients may be determined based on the signal model, the model parameters, and the analysis filter bank. In particular, the prediction coefficients may be determined based on an analytical evaluation of the signal model and the analysis filter bank. The analytical evaluation of the signal model and the analysis filter bank may lead to the determination of a lookup table and / or an analytical function. Thus, the prediction coefficients may be determined using the lookup table and / or the analytical function. Here, the lookup table and / or the analytical function may be predetermined based on the signal model and the analysis filter bank. The lookup table and / or the analytical function may provide the prediction coefficient(s) as a function of a parameter derived from the model parameter(s). The parameter derived from the model parameter(s) may, for example, be the model parameter(s) or may be obtained from the model parameter(s) using a predetermined function. Thus, the one or more prediction coefficients may be determined in a computationally efficient manner using a predetermined lookup table and / or analytical function that provides the one or more prediction coefficients depending only on the one or more parameters derived from the one or more model parameters. Thus, determining the prediction coefficients may simply be reduced to finding an entry in a lookup table. As indicated above, the analysis filter bank may have or exhibit a modulated structure. As a result of such a modulated structure, it is observed that the absolute values of the one or more prediction coefficients are independent of the index number of the first subband. This means that the lookup table and / or analytical function may be shift-invariant (except for the sign value) with respect to the index numbers of the subbands.In such cases, the parameters derived from the model parameters, i.e. the parameters input to a look-up table and / or analytical function to determine the prediction coefficients, may be derived by expressing the model parameters in a manner relative to a subband of the plurality of subbands.
[0018] As outlined above, the model parameters may indicate a fundamental frequency Ω of a multi-sinusoidal signal model (e.g., a periodic signal model). In such a case, determining the prediction coefficients may include determining a multiple of the fundamental frequency Ω that is within the first subband. If a multiple of the fundamental frequency Ω is within the first subband, a relative offset of the multiple of the fundamental frequency Ω from a center frequency of the first subband may be determined. In particular, the relative offset of the multiple of the fundamental frequency Ω that is closest to the center frequency of the first subband may be determined. A lookup table and / or analytical function may be predetermined such that the lookup table and / or analytical function provides the prediction coefficients as a function of possible relative offsets from a center frequency of a subband (e.g., as a function of normalized frequency f and / or as a function of a shift parameter Θ described herein). Thus, the prediction coefficients may be determined based on the lookup table and / or analytical function using the determined relative offsets. The predetermined lookup table may include a limited number of entries for a limited number of possible relative offsets. In such a case, before looking up the prediction coefficients from the look-up table, the determined relative offset may be rounded to the nearest possible relative offset from said limited number of possible relative offsets.
[0019] On the other hand, if no multiple of the fundamental frequency Ω is within the first subband, or rather within an extended frequency range around the first subband, the prediction coefficient may be set to 0. In such a case, the estimate of the first sample may also be 0.
[0020] Determining the prediction coefficients may include selecting one of a plurality of lookup tables based on the model parameters. For example, the model parameters may indicate a fundamental frequency Ω of a periodic signal model. The fundamental frequency Ω of the periodic signal model corresponds to a periodicity T of the periodic signal model. This document shows that for relatively small periodicities T, the periodic signal model converges toward a single sinusoidal model. This document also shows that for relatively large periodicities T, the lookup table varies slowly with the absolute value of T and depends primarily on the relative offset (i.e., the shift parameter Θ). Thus, multiple lookup tables may be predetermined for different values of the periodicity T. The model parameters (i.e., the periodicity T) may be used to select an appropriate one of the multiple lookup tables, and prediction coefficients may be determined (using the relative offset, e.g., the shift parameter Θ) based on the selected one of the multiple lookup tables. Thus, a model parameter (e.g., representing the periodicity T) that may have relatively high precision may be decoded into a pair of parameters (e.g., the periodicity T and the relative offset) of reduced precision, where the first parameter (e.g., the periodicity T) may be used to select a particular lookup table, and the second parameter (e.g., the relative offset) may be used to identify an entry within the selected lookup table.
[0021] The method may further include determining an estimate of the first sample by applying a prediction coefficient to a previous sample. Applying the prediction coefficient to the previous sample may include multiplying the prediction coefficient by a value of the previous sample, thereby obtaining an estimate of the first sample. Typically, the first samples of the first subband signal are determined by applying a prediction coefficient to a sequence of previous samples. Determining the estimate of the first sample may further include applying a scaling gain to the prediction coefficient and / or the first sample. The scaling gain (or an indicator thereof) may be used, for example, for long-term prediction (LTP). That is, the scaling gain may result from different predictors (e.g., long-term predictors). The scaling gain may be different for different subbands. Furthermore, the scaling gain may be transmitted as part of the encoded audio signal. Thus, using a signal model described by model parameters provides an efficient description of a subband predictor (including one or more prediction coefficients). The model parameters may be used to determine one or more prediction coefficients of a subband predictor. That is, an audio encoder does not need to transmit indicators of the one or more prediction coefficients, but only indicators of the model parameters. Typically, model parameters can be encoded more efficiently (i.e., using fewer bits) than the one or more prediction coefficients. Thus, the use of model-based prediction enables low-bitrate subband encoding.
[0022] The method may further include determining a prediction mask indicating multiple previous samples in multiple prediction mask support subbands. The multiple prediction mask support subbands may include at least one of the multiple subbands different from the first subband. Thus, the subband predictor may be configured to estimate samples of the first subband signal from samples of one or more other subband signals different from the first subband signal from the multiple subband signals. The prediction mask may define an arrangement of the multiple previous samples used to estimate the first sample of the first subband signal (e.g., a time lag with respect to the time slot of the first sample and / or a subband index lag with respect to the index number of the first subband).
[0023] The method may proceed by determining prediction coefficients to be applied to the previous samples. The prediction coefficients may be determined based on the signal model, the model parameters, and the analysis filter bank (e.g., using the model-based prediction schemes outlined above and herein). Thus, the prediction coefficients may be determined using one or more model parameters. In other words, a limited number of model parameters may be sufficient to determine the prediction coefficients. That is, by using model-based subband prediction, cross-subband prediction may be implemented in a bitrate-efficient manner.
[0024] The method may include determining an estimate of the first sample by applying the plurality of prediction coefficients to the plurality of previous samples, respectively, wherein determining the estimate of the first sample typically comprises determining a sum of the plurality of previous samples weighted by the plurality of respective prediction coefficients.
[0025] As outlined above, the model parameters may be indicative of a periodicity T. The plurality of lookup tables used to determine the one or more prediction coefficients may include lookup tables for different values of periodicity T. In particular, the plurality of lookup tables may be indicative of a function of [T min ,T max ]. As outlined in this paper, T min may be in the range of 0.25, and T max may be in the range of 2.5. min is T <T min may be chosen such that the audio signal can be modeled using a signal model having a single sinusoidal model component. max is T>T max For periodicity T max Or T max +1 lookup table is effectively periodic T max -1 to T max The same applies in general for n≧0, for periodicity T max +n or T max This is typically also true for +n+1.
[0026] The method may include determining the selected lookup table as the lookup table for the periodicity T indicated by the model parameter. After selecting a lookup table that includes or indicates the one or more prediction coefficients, the lookup parameter may be used to identify appropriate one or more entries in the selected lookup table, each of which indicates the one or more prediction coefficients. The lookup parameter may correspond to or be derived from a shift parameter Θ.
[0027] The method uses periodicity T>T maxFor model parameters that exhibit a residual periodicity T, subtract an integer value from T. r where the residual periodicity T r is [T max -1,T max ]. Then, the lookup table for determining the prediction coefficients is r may be determined as a lookup table for
[0028] The method uses periodicity T <T min For model parameters showing periodicity T min Further, a lookup parameter (e.g., a shift parameter Θ) for identifying the one or more entries of the selected lookup table that provides the one or more prediction coefficients may be selected based on the ratio T min / T. The one or more prediction coefficients may then be determined using the selected lookup table and the scaled lookup parameters. In particular, the one or more prediction coefficients may be determined based on the one or more entries of the selected lookup table that correspond to the scaled lookup parameters.
[0029] Therefore, the number of lookup tables is determined by a predetermined range [T min ,T max ], thereby limiting the memory requirements of the audio encoder / decoder. Nevertheless, the prediction coefficients may be determined for all possible values of the periodicity T using predetermined look-up tables, thereby allowing a computationally efficient implementation of the audio encoder / decoder.
[0030] According to a further aspect, a method for estimating a first sample of a first subband signal of an audio signal is described. As outlined above, the first subband of the audio signal may be determined using an analysis filter bank having a plurality of analysis filters providing a plurality of subband signals in a plurality of subbands from the audio signal. The above-mentioned features are also applicable to the following method.
[0031] The method may include determining a prediction mask indicative of a plurality of previous samples in a plurality of prediction mask support subbands, the plurality of prediction mask support subbands may include at least one of the plurality of subbands different from the first subband. In particular, the plurality of prediction mask support subbands may include the first subband and / or the plurality of prediction mask support subbands may include one or more of the plurality of subbands immediately adjacent to the first subband.
[0032] The method may further include determining a plurality of prediction coefficients to be applied to the plurality of previous samples. The plurality of previous samples are typically derived from the plurality of subband signals of the audio signal. In particular, the plurality of previous samples typically correspond to samples of a plurality of decoded subband signals. The plurality of prediction coefficients may correspond to prediction coefficients of a recursive (finite impulse response) prediction filter that also takes into account one or more samples of subbands different from the first subband. An estimate of the first sample may be determined by applying the plurality of prediction coefficients to the plurality of previous samples, respectively. Thus, the method enables subband prediction using one or more samples from other (e.g., adjacent) subbands. This may reduce aliasing artifacts caused by subband prediction-based coders.
[0033] The method may further include determining model parameters of a signal model. The prediction coefficients may be determined based on the signal model, the model parameters, and the analysis filter bank. Thus, the prediction coefficients may be determined using the model-based prediction described herein. In particular, the prediction coefficients may be determined using a lookup table and / or an analytical function. The lookup table and / or analytical function may be pre-determined based on the signal model and the analysis filter bank. Furthermore, the lookup table and / or analytical function may provide the prediction coefficients as a function of (only) parameters derived from the model parameters. Thus, the model parameters may directly provide the prediction coefficients using a lookup table and / or an analytical function. Thus, the model parameters may be used to efficiently describe the coefficients of a trans-subband predictor.
[0034] According to a further aspect, a method of encoding an audio signal is described. The method may include determining a plurality of subband signals from the audio signal using an analysis filter bank having a plurality of analysis filters. The method may proceed with estimating samples of the plurality of subband signals using any of the prediction methods described herein, thereby providing a plurality of estimated subband signals. Further, samples of a plurality of prediction error subband signals may be determined based on corresponding samples of the plurality of subband signals and samples of the plurality of estimated subband signals. The method may proceed with quantizing the plurality of prediction error subband signals and generating an encoded audio signal. The encoded audio signal may be indicative of (e.g., include) the plurality of quantized prediction error subband signals. Further, the encoded signal may be indicative of (e.g., include) one or more parameters used to estimate samples of the plurality of estimated subband signals. For example, it may be indicative of one or more model parameters used to determine one or more prediction coefficients subsequently used to estimate samples of the plurality of estimated subband signals.
[0035] According to another aspect, a method of decoding an encoded audio signal is described. The encoded audio signal typically indicates quantized prediction error subband signals and one or more parameters to be used to estimate samples of a plurality of estimated subband signals. The method may include dequantizing the plurality of quantized prediction error subband signals to provide a plurality of dequantized prediction error subband signals.
[0036] Further, the method may include estimating samples of the estimated subband signals using any of the prediction methods described herein. Samples of the decoded subband signals may be determined based on corresponding samples of the estimated subband signals and based on samples of the dequantized prediction error subband signals. A decoded audio signal may be determined from the decoded subband signals using a synthesis filter bank including a plurality of synthesis filters.
[0037] According to a further aspect, a system configured to estimate one or more first samples of a first subband signal of an audio signal is described. The first subband signal of the audio signal may be determined using a analysis filter bank having a plurality of analysis filters providing a plurality of subband signals from the audio signal in a plurality of respective subbands. The system may include a predictor calculator configured to determine model parameters of a signal model. Furthermore, the predictor calculator may be configured to determine one or more prediction coefficients to be applied to one or more previous samples of a first decoded subband signal derived from the first subband signal. Thus, the predictor calculator may be configured to determine one or more prediction coefficients of a recursive prediction filter, in particular a recursive subband prediction filter. The one or more prediction coefficients may be determined based on the signal model, the model parameters, and the analysis filter bank (e.g., using a model-based prediction method described herein). The one or more previous sample time slots are typically prior to the one or more first sample time slots. The system may further include a subband predictor configured to determine estimates of the one or more first samples by applying the one or more prediction coefficients to the one or more previous samples.
[0038] According to another aspect, a system configured to estimate one or more first samples of a first subband signal of an audio signal is described. The first subband signal corresponds to a first subband of a plurality of subbands. The first subband signal is typically determined using a analysis filter bank having a plurality of analysis filters that provide a plurality of subband signals for the plurality of subbands. The system includes a predictor calculator configured to determine a prediction mask indicative of a plurality of previous samples in a plurality of prediction mask support subbands, the plurality of prediction mask support subbands including at least one of the plurality of subbands different from the first subband. The predictor calculator is further configured to determine a plurality of prediction coefficients (or recursive prediction filters) to be applied to the plurality of previous samples. The system further includes a subband predictor configured to determine estimates of the one or more first samples by applying the plurality of prediction coefficients to the plurality of previous samples, respectively.
[0039] According to another aspect, an audio encoder configured to encode an audio signal is described. The audio encoder includes a analysis filter bank configured to determine a plurality of subband signals from the audio signal using a plurality of analysis filters. The audio encoder further includes a predictor calculator and a subband predictor as described herein configured to estimate samples of the plurality of subband signals, thereby providing a plurality of estimated subband signals. The encoder may further include a difference unit configured to determine a plurality of prediction error subband signal samples based on the plurality of subband signals and corresponding samples of the plurality of estimated subband signals. A quantization unit may be used to quantize the plurality of prediction error subband signals. A bitstream generation unit may further be configured to generate an encoded audio signal. The encoded audio signal indicates the plurality of quantized prediction error subband signals and one or more parameters (e.g., one or more model parameters) used to estimate samples of the plurality of estimated subband signals.
[0040] According to a further aspect, an audio decoder configured to decode an encoded audio signal is described. The encoded audio signal indicates (e.g., includes) the plurality of quantized prediction error subband signals and one or more parameters used to estimate samples of a plurality of estimated subband signals. The audio decoder may include an inverse quantizer configured to dequantize the plurality of quantized prediction error subband signals, thereby providing a plurality of dequantized prediction error subband signals. The decoder further includes a predictor calculator and a subband predictor described herein configured to estimate samples of the plurality of estimated subband signals. A summing unit may be used to determine samples of a plurality of decoded subband signals based on corresponding samples of the plurality of estimated subband signals and based on samples of the plurality of dequantized prediction error subband signals. Furthermore, a synthesis filter bank may be used to determine a decoded audio signal from the plurality of decoded subband signals using a plurality of synthesis filters.
[0041] According to a further aspect, a software program is described, the software program may be adapted for execution on a processor such that when executed on the processor it performs the method steps outlined herein.
[0042] According to another aspect, a storage medium is described that may include a software program for execution on a processor, the software program being adapted to perform the method steps outlined herein when executed on the processor.
[0043] According to a further aspect, a computer program product is described, which may include executable instructions for performing the method steps outlined herein when executed on a computer.
[0044] It should be noted that the methods and systems, including the preferred embodiments, outlined in this patent application may be used alone or in combination with other methods and systems disclosed herein. Furthermore, all aspects of the methods and systems outlined in this patent application may be combined in any manner. In particular, the features of the claims may be combined with each other in any manner. [Brief explanation of the drawings]
[0045] The present invention will now be described by way of illustrative examples, without limiting the scope or spirit of the invention, with reference to the accompanying drawings, in which: [Figure 1] 1 is a block diagram of an exemplary audio decoder that applies linear prediction in the filterbank domain (i.e., in the subband domain). [Figure 2] FIG. 1 illustrates an exemplary prediction mask in a time-frequency grid. [Figure 3] FIG. 10 shows exemplary tabulated data for a sinusoidal model-based predictor calculator. [Figure 4] FIG. 10 illustrates an exemplary noise shaping resulting from intra-band sub-band prediction. [Figure 5] FIG. 10 illustrates an exemplary noise shaping resulting from cross-band subband prediction. [Figure 6a] FIG. 1 depicts an exemplary two-dimensional quantization lattice underlying the tabulated data for a periodic model-based predictor calculation. [Figure 6b] FIG. 10 illustrates the use of different prediction masks for different ranges of signal periodicity. [Figure 7a] 1 is a flowchart of an exemplary encoding method using model-based subband prediction. [Figure 7b] 1 is a flowchart of an exemplary decoding method using model-based subband prediction. DETAILED DESCRIPTION OF THE INVENTION
[0046] The embodiments described below are merely illustrative of the principles of the present invention for model-based prediction in critically sampled filter banks. It is understood that modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. It is therefore the intention to be limited only by the scope of the appended claims and not by the specific details presented in the description and explanation of the embodiments herein.
[0047] 1 illustrates a block diagram of an exemplary audio decoder 100 that applies linear prediction in the filterbank domain (also referred to as the subband domain). The audio decoder 100 receives a bitstream containing information about a prediction error signal (also referred to as the residual signal) and possibly information about a predictor description used in a corresponding encoder to determine the prediction error signal from an original input audio signal. The information about the prediction error signal may relate to a subband of the input audio signal, and the information about the predictor description may relate to one or more subband predictors.
[0048] Given the received bitstream information, the inverse quantizer 101 may output samples 111 of a prediction error subband signal. These samples may be added to the output 112 of the subband predictor 103, and the sum 113 may be passed to the subband buffer 104. The subband buffer 114 maintains a record of previously decoded samples 113 of the subbands of the decoded audio signal. The output of the subband predictor 103 may be referred to as an estimated subband signal 112. The decoded samples of the subbands of the decoded audio signal may be submitted to the synthesis filter bank 102. The synthesis filter bank 102 transforms the subband samples into the time domain, thereby providing time-domain samples 114 of the decoded audio signal.
[0049] In other words, the decoder 100 may operate in the subband domain. In particular, the decoder 100 may use a subband predictor 103 to determine a plurality of estimated subband signals 112. Furthermore, the decoder 100 may use an inverse quantizer 101 to determine a plurality of residual subband signals 111. Each pair of the plurality of estimated subband signals 112 and the plurality of residual subband signals 111 may be summed to provide a corresponding plurality of decoded subband signals 113. The plurality of decoded subband signals 113 may be submitted to a synthesis filterbank 102 to provide a time-domain decoded audio signal 114.
[0050] In one embodiment of the subband predictor 103, a given sample of a given estimated subband signal 112 may be obtained by a linear combination of subband samples in the buffer 104 corresponding to a different time and frequency (i.e., a different subband) than the given sample of the given estimated subband signal 112. In other words, a sample of the estimated subband signal 112 in a first subband at a first time point may be determined based on one or more samples of the decoded subband signal 113 relating to a second time point (different from the first time point) and relating to a second subband (different from the first subband). The collection of prediction coefficients and their attachment to a time and frequency mask may define the predictor 103, and this information may be provided by the predictor calculator 105 of the decoder 100. The predictor calculator 105 outputs information defining the predictor 103 by transforming signal model data included in the received bitstream. An additional gain may be transmitted to modify the scaling of the output of the predictor 103. In one embodiment of the predictor calculator 105, the signal model data is provided in the form of an efficient parameterized line spectrum, where each line in the parameterized line spectrum or a series of lines in the parameterized line spectrum is used to point to tabulated values of predictor coefficients. Thus, the signal model data provided in the received bitstream may be used to identify entries in a predetermined lookup table. The entries from the lookup table provide one or more values for the predictor coefficients (also referred to as prediction coefficients) to be used by the predictor 103. The method applied for the table lookup may depend on a trade-off between complexity and memory requirements. For example, a nearest-neighbor type search may be used to achieve the lowest complexity, while an interpolation search method may provide similar performance with a smaller table size.
[0051] As indicated above, the received bitstream may include one or more explicitly signaled gains (or indicators of explicitly signaled gains). The gains may be applied as part of the predictor operation or after the predictor operation. The one or more explicitly signaled gains may be different for different subbands. The (indication of) the explicitly signaled additional gains is provided in addition to one or more model parameters used to determine the prediction coefficients of the predictor 103. Thus, the additional gains may be used to scale the prediction coefficients of the predictor 103.
[0052] Figure 2 shows an exemplary prediction mask support in a time-frequency lattice. Prediction mask support may be used for a predictor 103 operating on a filter bank with uniform time-frequency resolution, such as a cosine-modulated filter bank (e.g., an MDCT filter bank). The notation is illustrated by drawing 201, where the target dark-shaded subband samples 211 are the output of prediction based on the light-shaded subband samples 212. In drawings 202-205, the set of light-shaded subband samples represents the predictor mask support. The combination of source subband samples 212 and target subband samples 211 is referred to as prediction mask 201. A time-frequency lattice may be used to position subband samples in the vicinity of the target subband samples. The time slot index increases from left to right, and the subband frequency index increases from bottom to top. It should be noted that Figure 2 shows an exemplary case of prediction mask and predictor mask support; various other prediction masks and predictor mask supports may be used. An exemplary prediction mask is as follows:
[0053] The prediction mask 202 defines an intra-band prediction of an estimated subband sample 221 at time instant k from two previously decoded subband samples 222 at time instants k-1 and k-2.
[0054] The prediction mask 203 defines a cross-band prediction of an estimated subband sample 231 in subband n at time instant k based on three previously decoded subband samples 232 in subbands n-1, n, n+1 at time instant k-1.
[0055] The prediction mask 204 defines a cross-band prediction of three estimated subband samples 241 in three different subbands n-1, n, n+1 at time point k based on three previously decoded subband samples 242 in subbands n-1, n, n+1 at time point k-1. The cross-band prediction may be performed such that each estimated subband sample 241 can be determined based on all three previously decoded subband samples 242 in subbands n-1, n, n+1.
[0056] The prediction mask 205 defines a cross-band prediction of an estimated subband sample 251 in subband n at time instant k based on 12 previously decoded subband samples 252 in subbands n-1, n, n+1 at time instants k-2, k-3, k-4, k-5.
[0057] FIG. 3 shows tabulated data for the sinusoidal model-based predictor calculator 105 operating on a cosine-modulated filter bank. The prediction mask support is from plot 204. For a given frequency parameter, the subband with the closest subband center frequency may be selected as the center target subband. The difference between the frequency parameter and the center frequency of the center target subband may be calculated in units of the filter bank's frequency interval (bin). This gives a value between -0.5 and 0.5, which may be rounded to the nearest available entry in the tabulated data and is depicted by the horizontal axis of the nine graphs 301 in FIG. 3. This generates a 3×3 matrix of coefficients to be applied to the most recent values of the multiple decoded subband signals 113 in the subband buffer 104 for the target subband and its two adjacent subbands. The resulting 3×1 vector constitutes the contribution of the subband predictor 103 to these three subbands for the given frequency parameter. The process may be repeated in an additive manner for all sinusoidal components in the signal model.
[0058] In other words, Figure 3 shows an example of a model-based description of a subband predictor. Suppose the input audio signal has fundamental frequencies Ω, Ω, …, Ω. M-13 to determine a 3×3 matrix of prediction coefficients by determining coefficient values 302 for relative frequency values 303 of the fundamental frequency Ω. This means that the coefficients for the subband predictor 103 using the prediction mask 204 can be determined using only the received information about the particular fundamental frequency Ω. In other words, by modelling the input audio signal using, for example, one model of multiple sinusoidal components, a bitrate-efficient description of the subband predictor can be provided.
[0059] Figure 4 shows an example noise shape resulting from intraband subband prediction in a cosine-modulated filter bank. The signal model used to perform intraband subband prediction is a second-order autoregressive stochastic process with peaked resonances, as described by a second-order differential equation driven by random Gaussian white noise. Curve 401 shows the measured magnitude spectrum for this process realization. For this example, the prediction mask 202 from Figure 2 is applied. That is, the predictor calculator 105 generates a subband predictor 103 for a given target subband 221 based on previous subband samples 222 only within the same subband. Replacing the inverse quantizer 101 with a Gaussian white noise generator results in the synthesized magnitude spectrum 402. As can be seen, strong aliasing artifacts appear in this synthesis, as the synthesized spectrum 402 has peaks that do not match those of the original spectrum 401.
[0060] FIG. 5 shows an exemplary noise shaping resulting from cross-band subband prediction. The setup is the same as in FIG. 4, except for the fact that a prediction mask 203 is applied. Thus, the calculator 105 generates a subband predictor 103 for a given target subband 231 based on previous subband samples 232 in the target subband and its two adjacent subbands. As can be seen from FIG. 5, the spectrum 502 of the synthesized signal substantially matches the spectrum 501 of the original signal. That is, when using cross-band subband prediction, aliasing problems are substantially suppressed. Thus, FIGS. 4 and 5 show that aliasing artifacts caused by subband prediction can be reduced when using cross-band subband prediction, i.e., when predicting a subband sample based on previous subband samples in one or more adjacent subbands. As a result, subband prediction can be applied even in the context of low-bitrate audio encoders without the risk of audible aliasing artifacts. While the use of cross-band subband prediction typically increases the number of prediction coefficients, as shown in the context of Figure 3, the use of a model for the input audio signal (e.g., the use of a sinusoidal or periodic model) allows for an efficient description of the subband predictor, thereby enabling the use of cross-band subband prediction for low bitrate audio coders.
[0061] In the following, a description of the principles of model-based prediction in critically sampled filter banks will be outlined in appropriate mathematical terminology with reference to Figures 1-6.
[0062] A possible signal model underlying linear prediction is that of a zero-mean weakly stationary stochastic process x(t), whose statistics are determined by the autocorrelation function r(τ) = E{x(t)x(t-τ)}. A good model for the critically sampled filter banks considered here is the α :α∈A} is a real-valued composite waveform w that forms an orthonormal basis. α(t). In other words, the filter bank is a set of waveforms {w α :α∈A}. The subband samples of the time domain signal s(t) can be expressed as the dot product
number
number
[0063]
number
[0064]
number
[0065]
number
[0066]
number
[0067]
number
[0068] In this paper, the predictor calculator 105 calculates the prediction coefficients {c β It is proposed to transmit a parametric representation of the signal model from which {\displaystyle \beta ∈ B} can be derived. For example, the signal model may provide a parametric representation of the autocorrelation function r(τ) of the signal model. The decoder 100 may use the received parametric representation to derive the autocorrelation function r(τ) and then convert the autocorrelation function r(τ) to the composite waveform cross-correlation W to derive the elements of the covariance matrix required for normal equation (7). αβ (τ) may be combined with (τ). The equation can then be solved to obtain the prediction coefficients.
[0069] In other words, the input audio signal to be encoded may be modeled by a process x(t) that can be described using a limited number of model parameters. In particular, the modeled process x(t) may be such that its autocorrelation function r(τ) = E{x(t)x(t-τ)} can be described using a limited number of parameters. The limited number of parameters describing the autocorrelation function r(τ) may be transmitted to the decoder 100. The predictor calculator 105 of the decoder 100 may determine the autocorrelation function r(τ) from the received parameters, and the covariance matrix R of the subband signals from which the normal equation (7) can be determined. αβ Equation (3) may then be used to determine the normal equation (7). The normal equation (7) may then be solved by the predictor calculator 105, thereby determining the prediction coefficients c β Give.
[0070] In the following, exemplary signal models are described that can be used to apply the above model-based prediction schemes in an efficient manner. The signal models described below are typically very important for coding audio signals, e.g., for coding speech signals.
[0071] An example of a signal model is a sinusoidal process
number
number
[0072] A generalization of such a sinusoidal process is a multi-sinusoidal model with a set of (angular) frequencies S, i.e. with multiple distinct (angular) frequencies ξ.
[0073]
number
[0074]
number
[0075]
number
number
[0076] An example of a compact description of the set S of frequencies for a multi-sine model is: 1. Single fundamental frequency Ω:S={Ων:ν=1,2,…} 2. M fundamental frequencies Ω0, Ω1, …, Ω M-1 :S={Ω k ν:ν=1,2,…, k=0,1,…,M-1} 3. Single-sided band-shifted fundamental frequencies Ω,θ:S={Ω(ν+θ):ν=1,2,…} 4. Slightly anharmonic model: Ω, a: S = {Ων (1 + aν 2 ) 1 / 2 : ν=1,2,…} where a describes the anharmonic component of the model.
[0077] Thus, a (possibly relaxed) multi-sine model representing the PSD given by equation (12) can be described in an efficient manner using one of the exemplary descriptions listed above. For example, the complete set of frequencies S of the line spectrum of equation (12) may be described using only a single fundamental frequency Ω. If the input audio signal to be encoded can be well described using a multi-sine model representing a single fundamental frequency Ω, the model-based predictor may be described by a single parameter (i.e., by the fundamental frequency Ω) regardless of the number of prediction coefficients used by the subband predictor 103 (i.e., regardless of the prediction masks 202, 203, 204, 205).
[0078] Case 1 describes a set of frequencies S, giving a process x(t) that models an input audio signal with period T = 2π / Ω. If a zero-frequency (DC) contribution with variance 1 / 2 is included in equation (11) and the result is rescaled by a factor 2 / T, the autocorrelation function of the periodic model process x(t) becomes
number
[0079] Using the relaxation factor definition ρ=exp(−Tε), the autocorrelation function of the relaxed version of the periodic model is given by:
[0080]
number
number
[0081] The global signal model described above typically includes a sinusoidal amplitude parameter a ξ ,b ξDue to the unity variance assumption, it has a flat large-scale power spectrum. However, signal models are typically only considered locally for a subset of the subbands of a critically sampled filter bank, where the filter bank is responsible for shaping the entire spectrum. In other words, for signals with spectral shapes that vary slowly compared to the subband width, a flat power spectrum model provides a good match to the signal, and thus model-based predictors offer a sufficient level of prediction gain. More generally, PSD models can be described using standard parameterizations of autoregressive (AR) or autoregressive moving average (ARMA) processes. This will improve the performance of model-based predictions, possibly at the expense of increasing the descriptive model parameters.
[0082] Another variant is obtained by abandoning the stationary assumption for the stochastic signal model. In that case, the autocorrelation function becomes a function of two variables: r(t,s) = E{x(t)x(s)}. For example, the relevant non-stationary sinusoidal model may include amplitude modulation (AM) and frequency modulation (FM).
[0083] Furthermore, more deterministic signals may be used. As will be seen in some of the examples below, predictions may have errors that go to zero in some cases. In such cases, a probabilistic approach can be avoided. When predictions are perfect for all signals in the model space, there is no need to average the prediction performance by a probability measure over the model space considered.
[0084] In the following, various aspects of modulated filter banks are described. In particular, those aspects that affect the determination of the covariance matrix and thereby provide an efficient means for determining the prediction coefficients of the subband predictors are described. A modulated filter bank may be described as a two-dimensional index set of composite waveforms, α = (n, k), where n = 0, 1, ... are subband indices (frequency bands) and k ∈ Z are subband sample indices (time slots). For ease of presentation, the composite waveforms are assumed to be given in continuous time and normalized to a unit time stride.
[0085] w n,k (t)=u n (tk) (16) where, for a cosine modulated filter bank,
number
[0086] Due to the shift-invariant structure, it can be seen that the cross-correlation function of the composite waveform (defined in equation (4)) can be written as:
[0087]
number
[0088]
number
[0089]
number
[0090]
number
[0091] For a given signal model autocorrelation function r(τ), the formula above can be inserted into the definition of the subband sample covariance matrix given by equation (3):
number
[0092] As a function of the power spectral density P(ω) of a given signal model (which corresponds to the Fourier transform of the autocorrelation function r(τ)),
number
[0093]
number
[0094] Equation (24) provides an efficient means of determining the coefficients of the subband sample covariance matrix when knowing the PSD of the underlying signal model. As an example, for a sinusoidal model-based prediction scheme utilizing a signal model x(t) containing a single sinusoid at (angular) frequency ξ, the PSD is given by P(ω) = (1 / 2)(δ(ω-ξ) + δ(ω+ξ)). Substituting P(ω) into Equation (24) gives four terms, three of which can be ignored under the assumption that n+m+1 is large. The remaining terms are:
[0095]
number
[0096]
number
[0097] In the following, the solution of normal equation (26) for various prediction mask supports (as shown in FIG. 2) is given in an exemplary manner. An example of a causal second-order intra-band predictor is obtained by choosing the prediction mask support B={(p,-1),(p,-2)}. This prediction mask support corresponds to the prediction mask 202 in FIG. 2. Using the approximation of equation (25), normal equation (26) for this 2-tap prediction becomes:
[0098]
number
number
[0099] As discussed above with respect to Figure 4, intra-band prediction has certain drawbacks related to aliasing artifacts in noise shaping. The next example concerns the improved behavior shown in Figure 5. Causal cross-band prediction as taught in this paper is obtained by choosing the prediction mask support B = {(p-1, -1), (p, -1), (p+1, -1)}. This requires only one preceding time slot instead of two, and produces noise shaping with less alias frequency contributions than the classical prediction mask 202 in the first example. The prediction mask support B = {(p-1, -1), (p, -1), (p+1, -1)} corresponds to the prediction mask 203 in Figure 2. The normal equation (26), based on the approximation of Equation (25), is then calculated using the three unknown coefficients c m This reduces to two equations for [-1], m=p-1, p, p+1.
[0100]
number
[0101]
number
[0102] As shown by the prediction mask 204 in FIG. 2, three subband samples 〈x, w m,0 By using the same prediction mask support B={(p-1,-1),(p,-1),(p+1,-1)} to predict p, we obtain a 3x3 prediction matrix. Introducing a more natural strategy to avoid ambiguity in the normal equations, i.e.,
number
[0103] Thus, it has been shown that a signal model x(t) may be used to describe the underlying characteristics of the input audio signal to be encoded. Parameters describing the autocorrelation function r(τ) may be transmitted to the decoder 100, thereby enabling the decoder 100 to compute a predictor from the transmitted parameters and knowledge of the signal model x(t). For modulated filter banks, it has been shown that efficient means can be derived for determining the subband covariance matrix of the signal model and for solving the normal equations to determine the predictor coefficients. In particular, it has been shown that the resulting predictor coefficients are invariant to subband shifts and typically depend only on the normalized frequency for a particular subband. As a result, a predetermined lookup table (e.g., as shown in FIG. 3) can be provided that allows the determination of the predictor coefficients knowing the normalized frequency f, which is independent (except for the parity value) of the subband index p at which the predictor coefficients are determined.
[0104] In the following, periodic model-based predictions, for example using a single fundamental frequency Ω, are described in more detail. The autocorrelation function r(τ) of such a periodic model is given by equation (13). The equivalent PSD or line spectrum is
number
[0105] When the period T of the periodic model is small enough, say T≦1, the fundamental frequency Ω=2π / T is large enough to allow the application of a sinusoidal model as derived above, with the partial frequency ξ=qΩ closest to the center frequency π(p+(1 / 2)) of the subband p of the target subband sample to be predicted. This means that periodic signals with small periods T, i.e., small relative to the time stride of the filter bank, can be well modeled and predicted using the sinusoidal model above.
[0106] When the period T is sufficiently large compared to the duration K of the filter bank window v(t), the predictor reduces to an approximation of the delay by T. As will be shown, the coefficients of this predictor can be read directly from the waveform cross-correlation function given by equation (19).
[0107] Substituting the model based on equation (13) into equation (22) gives:
[0108]
number
[0109]
number
number
[0110]
number
number
[0111] In view of the above, we propose a method to estimate the subband sample 〈x,w〉 (from subband p and time index 0) by using a suitable prediction mask support B centered at (p, -T) with a time diameter approximately equal to T. p,0 It is taught that > can be predicted. The normal equations may be solved for each value of T and p. In other words, for each periodicity T of the input audio signal, for each subband p, the prediction coefficients for a given prediction mask support B may be determined using the normal equations (33).
[0112] Given the large number of subbands p and the wide range of periods T, it is impractical to directly tabulate all predictor coefficients. However, similar to the sinusoidal model, the modulated structure of the filter bank offers a significant reduction in the required table size through its invariance with respect to frequency shifts. Typically, it will suffice to examine a shifted harmonic model with shift parameter -1 / 2<θ≦1 / 2 around the center of subband p, i.e., centered around π(p+(1 / 2)), defined by a subset S(θ) of positive frequencies from the set of frequencies π(p+(1 / 2))+(q+θ)Ω, q∈Z.
[0113]
number
[0114]
number
number
[0115] Equation (36) expresses the prediction coefficients c for subband (p+ν) at time index k. p+ν [k], where the sample to be predicted is the sample from subband p at time index 0. As can be seen from equation (36), the prediction coefficients c p+ν [k] is the target subband index p, and is a factor that affects the sign of the prediction coefficient (-1). pk However, the absolute value of the prediction coefficient is independent of the target subband index p. On the other hand, the prediction coefficient c p+ν [k] depends on the periodicity T and the shift parameter θ. Furthermore, the prediction coefficient c p+ν[k] depends on v and k, i.e., the prediction mask support B used to predict the target sample in the target subband p.
[0116] In this paper, we consider a set of prediction coefficients c p+ν It is proposed to provide a lookup table that allows searching the prediction coefficients c for a predetermined set of values of the period T and the shift parameter θ. p+ν The set [k] is given. To limit the number of lookup table entries, the number of predetermined values of the period T and the number of predetermined values of the shift parameter θ should be limited. As can be seen from Equation (36), the preferred quantization step size for the predetermined values of the periodicity T and the shift parameter θ should depend on the periodicity T. In particular, for a relatively large periodicity T (relative to the window function duration K), a relatively large quantization step for the periodicity T and the shift parameter θ may be used. Conversely, for a relatively small periodicity T approaching 0, only one sinusoidal contribution needs to be taken into account, and thus the periodicity T loses its importance. On the other hand, since the formula for sinusoidal prediction based on Equation (29) requires that the normalized absolute frequency shift f = Ωθ / π = (1 / 2)θ / T varies slowly, the quantization step size for the shift parameter θ should be scaled based on the periodicity T.
[0117] Overall, this paper proposes using uniform quantization of the periodicity T with a fixed step size. However, the shift parameter θ may also be quantized in a uniform manner with a step size proportional to min(T,A), where the value of A depends on the details of the filter window function. Furthermore, for T<2, the range of the shift parameter θ may be restricted to |θ|≦min(CT,1 / 2) for some constant C, reflecting a restriction on the absolute frequency shift f.
[0118] Figure 6a shows an example of the resulting quantization lattice in the (T, θ) plane for A = 2. Only in the intermediate range of 0.25 ≤ T ≤ 1.5 is the full two-dimensional dependence considered. Meanwhile, for the remaining range of interest, an essentially one-dimensional parameterization such as that given by Eqs. (29) and (36) can be used. In particular, for periodicities T approaching 0 (e.g., T < 0.25), the periodic model-based prediction essentially corresponds to the sinusoidal model-based prediction, and the prediction coefficients may be determined using formula (29). On the other hand, for periodicities T substantially exceeding the window duration K (e.g., T > 1.5), the prediction coefficients c using the periodic model-based prediction are p+ν The set of [k] may be determined using equation (36), which can be reinterpreted by the substitution θ=φ+(1 / 4)Tν.
[0119]
number
[0120] The modified offset parameter φ can be interpreted as a shift in the harmonic series in units of the fundamental frequency, measured from the midpoint of the midpoints of the source and target bins. It is advantageous to maintain this modified parameterization (T,φ) for all values of the periodicity T, since the symmetry in equation (37) that is evident for simultaneous sign changes of φ and ν holds in general and can be exploited to reduce table size.
[0121] As noted above, Figure 6a illustrates the two-dimensional quantization lattice underlying the tabulated data for a periodic model-based predictor computation in a cosine-modulated filter bank. The signal model is of a signal with period T 602, measured in units of the filter bank time step. Equivalently, the model includes frequency lines at integer multiples, also known as partials, of the fundamental frequency corresponding to period T. For each target subband, the shift parameter θ 601 indicates the distance to the center frequency of the nearest partial, measured in units of the fundamental frequency Ω. The shift parameter θ 601 has values between -0.5 and 0.5. The black crosses 603 in Figure 6a indicate the appropriate density of quantization points to tabulate a predictor with high prediction gain based on a periodic model. For large periods T (e.g., T > 2), the lattice is uniform. As the period T decreases, increased density in the shift parameter θ is typically required. However, in the region outside the line 604, the distance θ is greater than one frequency bin of the filter bank. Thus, most grid points in this region can be ignored. Polygon 605 defines the region that is sufficient for complete tabulation. In addition to the slightly outer slanted line of line 604, boundaries at T=0.25 and T=1.5 are introduced. This is made possible by the fact that small periods 602 can be treated as separate sinusoids, and the predictor for large periods 602 can be approximated by an essentially one-dimensional table that depends primarily on the shift parameter θ (or a modified shift parameter φ). For the embodiment shown in FIG. 6a, the prediction mask support is typically similar to the prediction mask 205 of FIG. 2 for large periods T.
[0122] Figure 6b shows the periodic model-based prediction for a relatively large period T and a relatively small period T. From the figure, it can be seen that for a large period T, i.e., for a relatively small fundamental frequency Ω 613, the filter bank's window function 612 captures a relatively large number of lines of the PSD of the periodic signal, or Dirac pulse 616. The Dirac pulse 616 is located at frequency ω = qΩ, where q ∈ Z. The center frequency of the filter bank's subband is located at frequency ω = π(p + (1 / 2)), where p ∈ Z. For a given subband p, the frequency location of the pulse 616 with frequency ω = qΩ closest to the given subband's center frequency ω = π(p + (1 / 2)) can be described in relative form as qΩ = π(p + (1 / 2)) + ΘΩ, with a shift parameter Θ ranging from -0.5 to +0.5. Thus, the term ΘΩ reflects the distance (in frequency) from the center frequency ω=π(p+(1 / 2)) to the nearest frequency component 616 of the harmonic model. This is shown in the upper diagram of Figure 6b, where the center frequency 617 is ω=π(p+(1 / 2)) and the distance 618 ΘΩ is shown for a relatively large period T. It can be seen that the shift parameter Θ allows us to describe the entire harmonic series from the point of view of the center of subband p.
[0123] The bottom diagram of Figure 6b shows the case for a relatively small period T, i.e., for a relatively large fundamental frequency Ω 623, in particular a fundamental frequency 623 larger than the width of the window 612. It can be seen that in such a case, the window function 612 may only contain a single pulse 626 of the periodic signal, which may therefore be viewed as a sinusoidal signal within the window 612. This means that for a relatively small period T, the periodic model-based prediction scheme converges towards a sinusoidal model-based prediction scheme.
[0124] 6b also shows exemplary prediction masks 611 and 621 that may be used for the periodic model-based prediction scheme and the sinusoidal model-based prediction scheme, respectively. The prediction mask 611 used for the periodic model-based prediction scheme may correspond to the prediction mask 205 of FIG. 2 and may include a prediction mask support 614 for estimating target subband samples 615. The prediction mask 621 used for the sinusoidal model-based prediction scheme may correspond to the prediction mask 203 of FIG. 2 and may include a prediction mask support 624 for estimating target subband samples 625.
[0125] FIG. 7a illustrates an exemplary encoding method 700 involving model-based subband prediction using a periodic model (e.g., including a single fundamental frequency Ω). A frame of an input audio signal is considered. For this frame, a periodicity T or a fundamental frequency Ω may be determined (step 701). The audio encoder may include elements of the decoder 100 shown in FIG. 1. In particular, the audio encoder may include a predictor calculator 105 and a subband predictor 103. The periodicity T or the fundamental frequency Ω may be determined such that the mean value of the squared prediction error subband signal 111 based on Equation (6) is reduced (e.g., minimized). As an example, the audio encoder may determine prediction error subband signals 111 using various fundamental frequencies Ω and apply a brute force approach to determine the fundamental frequency Ω such that the mean value of the squared prediction error subband signal 111 is reduced (e.g., minimized). The method proceeds by quantizing the resulting prediction error subband signal 111 (step 702). Furthermore, the method includes a step 703 of generating a bitstream including information indicative of the determined fundamental frequency Ω and the quantized prediction error sub-band signal 111 .
[0126] When determining the fundamental frequency Ω in step 701, the audio encoder may utilize equations (36) and / or (29) to determine prediction coefficients for the particular fundamental frequency Ω. The set of possible fundamental frequencies Ω may be limited by the number of bits available for transmission of information indicating the determined fundamental frequency Ω.
[0127] It should be noted that the audio coding system may use a predetermined model (e.g., a periodic model with a single fundamental frequency Ω or any of the other models given herein) and / or predetermined prediction masks 202, 203, 204, 205. On the other hand, the audio coding system may be given additional freedom by allowing the audio encoder to determine an appropriate model and / or appropriate prediction mask for the audio signal to be encoded. Information about the selected model and / or the selected prediction mask is then encoded into a bitstream and provided to the corresponding decoder 100.
[0128] FIG. 7b shows an example method 710 for decoding an audio signal encoded using model-based prediction. It is assumed that the decoder knows the signal model and prediction mask used by the encoder (either via the received bitstream or due to predetermined settings). Furthermore, for purposes of illustration, it is assumed that a periodic prediction model was used. The decoder 100 extracts information about the fundamental frequency Ω from the received bitstream (step 711). Using the information about the fundamental frequency Ω, the decoder 100 may determine the periodicity T. The fundamental frequency Ω and / or the periodicity T may be used to determine sets of prediction coefficients for various subband predictors (step 712). The subband predictors may be used to determine estimated subband signals (step 713), which are combined with the dequantized prediction error subband signals 111 to provide the decoded subband signals 113 (step 714). The decoded subband signals 113 are filtered using the synthesis filterbank 102 (step 715 ), thereby providing a decoded time-domain audio signal 114 .
[0129] The predictor calculator 105 may use equations (36) and / or (29) to determine prediction coefficients for the subband predictor 103 based on the received information about the fundamental frequency Ω (step 712). This may be performed in an efficient manner using lookup tables such as those shown in FIGS. 6a and 3. As an example, the predictor calculator 105 may determine the periodicity T and determine whether the periodicity is below a predetermined lower threshold (e.g., T = 0.25). If so, a sinusoidal model-based prediction scheme is used. That is, based on the received fundamental frequency Ω, a subband p having a multiple ω = qΩ of the fundamental frequency, where q ∈ Z, is determined. Then, a normalized frequency f is determined using the relationship ξ = π(p + (1 / 2) + f), where frequency ξ corresponds to the multiple ω = qΩ within subband p. Predictor calculator 105 may then use equation (29) or a pre-calculated look-up table to determine the set of prediction coefficients (e.g., using prediction mask 203 of FIG. 2 or prediction mask 621 of FIG. 6b).
[0130] It should be noted that a different set of prediction coefficients may be determined for each subband. However, in the case of a sinusoidal model-based prediction scheme, a set of prediction coefficients is typically determined only for the subband p that is significantly affected by the fundamental frequency multiple ω = qΩ, where q ∈ Z. For other subbands, no prediction coefficients are determined. That is, the estimated subband signal 112 for such other subbands is zero. To reduce the computational complexity of the decoder 100 (and the encoder using the same predictor calculator 105), the predictor calculator 105 may utilize a predetermined lookup table that provides a set of prediction coefficients depending on values for T and Θ. In particular, the predictor calculator 105 may utilize multiple lookup tables for multiple different values of T, each of which provides a different set of prediction coefficients for multiple different values of the shift parameter Θ.
[0131] In a practical implementation, multiple lookup tables may be provided for various values of the periodic parameter T. By way of example, lookup tables may be provided for values of T ranging from 0.25 to 2.5 (as shown in FIG. 6a). Lookup tables may be provided for different predetermined granularities or step sizes of the periodic parameter T. In one exemplary implementation, the step size for the normalized periodic parameter T is 1 / 16, and different lookup tables for quantized prediction coefficients are provided for T=8 / 32 to T=80 / 32. Thus, a total of 37 different lookup tables may be provided. Each table may provide quantized prediction coefficients as a function of the shift parameter Θ or as a function of the modified shift parameter φ.
[0132] For a range increased by half the step size, i.e., [9 / 32, 81 / 32], a lookup table for T=8 / 32 to T=80 / 32 may be used. For a given periodicity different from the available periodicity for which the lookup table is defined, the lookup table for the closest available periodicity may be used. As outlined above, for long periods T (e.g., for periods T exceeding the period for which the lookup table is defined), equation (36) may be used. Alternatively, for periods T exceeding the period for which the lookup table is defined, e.g., for periods T>81 / 32, the period T is an integer delay T i and residual delay T r where T = T i +T r The separation is performed by subtracting the residual delay T r However, equation (36) is applicable and a look-up table is available, e.g., in the interval [1.5, 2.5] or [49 / 32, 81 / 32] for the above example. By doing so, the prediction coefficients are calculated based on the residual delay T r , and the subband predictor 103 can be determined using a lookup table for integer delay T iIt may act on the sub-band buffer 104 delayed by. For example, if the period is T = 3.7, the integer delay is T i = 2, and the residual delay T r = 1.7 may follow. The predictor may be applied based on the coefficient for T r = 1.7 on the signal buffer. The signal buffer is delayed by (additional) T i = 2.
[0133] The above separation approach relies on the reasonable assumption that the extractor approximates the delay by T within the range [1.5, 2.5] or [49 / 32, 81 / 32]. The advantage of this separation procedure compared to using Equation (36) is that the prediction coefficients can be determined based on a computationally efficient table lookup operation.
[0134] As outlined above, for short periods (T < 0.25), Equation (29) may be used to determine the prediction coefficients. Alternatively, it may be beneficial to utilize a (already available) lookup table (to reduce the computational load). It is observed that the modified shift parameter φ is restricted to the range |φ| < T using a sampling pitch size of Δφ = T / 32 (for T < 0.25 and C = 1, A = 1 / 2).
[0135] In this paper, it is proposed to reuse the lookup table for the lowest period T = 0.25 by scaling the modified shift parameter φ with T l / T. Here, T l corresponds to the lowest period for which the lookup table is available (e.g., T l= 0.25). As an example, using T = 0.1 and φ = 0.07, the table for T = 0.25 may be queried using a rescaled shift parameter φ = (0.25 / 0.1) 0.07 = 0.175. By doing so, prediction coefficients for short periods (e.g., T < 0.25) can also be determined in a computationally efficient manner using a table lookup operation. Furthermore, since the number of lookup tables can be reduced, the memory requirements for the predictor can be alleviated.
[0136] In this paper, a model-based subband prediction scheme is described. The model-based subband prediction scheme allows for an efficient description of the subband predictor, i.e., a description that requires only a relatively small number of bits. As a result of the efficient description of the subband predictor, a cross-subband prediction scheme can be used, which leads to a reduction in aliasing artifacts. Overall, this allows for the implementation of low-bitrate audio coders using subband prediction.
[0137] Several aspects will be described. [Aspect 1] A method for estimating a first sample (615) of a first subband signal in a first subband of an audio signal, the first subband signal of the audio signal being determined using an analysis filter bank (612) having a plurality of analysis filters providing a plurality of subband signals in a plurality of subbands from the audio signal, respectively, the method comprising: determining model parameters (613) of the signal model; determining, based on the signal model, based on the model parameters (613), and based on the analysis filter bank (612), prediction coefficients to be applied to a previous sample (614) of a first decoded subband signal derived from the first subband signal, wherein the time slot of the previous sample (614) precedes the time slot of the first sample (615); determining an estimate of the first sample (615) by applying the prediction coefficients to the previous sample (614), method. [Aspect 2] the signal model includes one or more sinusoidal model components; the model parameters indicate frequencies of the one or more sinusoidal model components; The method of embodiment 1. Aspect 3 3. The method of aspect 2, wherein the model parameter indicates a fundamental frequency Ω of a multi-sinusoidal signal model. Aspect 4 the multi-sinusoidal signal model has a periodic signal component; the periodic signal component includes a plurality of sinusoidal components; the plurality of sinusoidal components have frequencies that are multiples of the fundamental frequency Ω; The method of embodiment 3. Aspect 5 5. The method of any one of aspects 1 to 4, wherein the method includes a plurality of model parameters (613) of the signal model. Aspect 6 the signal model comprises a plurality of periodic signal components; the plurality of model parameters are a plurality of fundamental frequencies Ω0, Ω1, ..., Ω of the plurality of periodic signal components; M-1 Showing, The method of embodiment 5. Aspect 7 7. The method of embodiment 5 or 6, wherein one or more of the plurality of signal model parameters indicates a shift and / or deviation of the signal model from a periodic signal model. Aspect 8 8. The method of any one of aspects 1-7, wherein determining the model parameters comprises extracting the model parameters from a received bitstream indicative of the model parameters and a prediction error signal. Aspect 9 determining the model parameters includes determining the model parameters such that a mean value of a squared prediction error signal is small; the prediction error signal is determined based on a difference between the first sample and the estimate of the first sample; 8. The method of any one of embodiments 1 to 7. Aspect 10 10. The method of aspect 9, wherein the average value of the squared prediction error signal is determined based on a plurality of consecutive first samples of the first subband signal. Aspect 11 the step of determining the prediction coefficients includes determining the prediction coefficients using a look-up table or an analytical function; the look-up table or the analytical function provides the prediction coefficients as a function of certain parameters derived from the model parameters; the lookup table or the analysis function is predetermined based on the signal model and based on the analysis filter bank; 11. The method of any one of embodiments 1 to 10. Aspect 12 the analysis filter bank has a modulated structure, The absolute value of the prediction coefficient is independent of the index number of the first subband; The method of embodiment 11. Aspect 13 the model parameter indicates a fundamental frequency Ω of a multi-sinusoidal signal model; the step of determining the prediction coefficients comprises determining multiples of the fundamental frequency Ω that are within the first subband; 13. The method of embodiment 11 or 12. Aspect 14 determining the prediction coefficients If the multiple of the fundamental frequency Ω is within the first subband, determining a relative offset of the multiple of the fundamental frequency Ω from a center frequency of the first subband; setting the prediction coefficients to zero if no multiple of the fundamental frequency Ω is within the first subband; The method of embodiment 13. Aspect 15 the look-up table or the analytical function gives the prediction coefficients as a function of possible relative offsets from the center frequency of a subband, determining the prediction coefficients includes determining the prediction coefficients based on the look-up table or the analytical function using the determined relative offsets; The method of embodiment 14. Aspect 16 the lookup table contains a limited number of entries for a limited number of possible relative offsets; the step of determining the prediction coefficients includes rounding the determined relative offset to the nearest possible relative offset from the limited number of possible relative offsets. The method of embodiment 15. Aspect 17 The step of determining the prediction coefficients comprises: selecting one of a plurality of lookup tables based on the model parameters; determining the prediction coefficients based on a selected one of the plurality of lookup tables; 17. The method of any one of embodiments 13 to 16. Aspect 18 said model parameters exhibit a periodicity T, the plurality of lookup tables include lookup tables for different values of periodicity T; the method includes determining the selected lookup table as the lookup table for a periodicity T indicated by the model parameters; The method of embodiment 17. Aspect 19 The plurality of lookup tables are min ,T max], and a look-up table for different values of periodicity T at a predetermined step size ΔT within the range T min is T <T min wherein the audio signal is such that it can be modeled using a signal model having a single sinusoidal model component; and / or T max is T>T max For periodicity T max Or T max +1 lookup table is effectively periodic T max -1 to T max , which corresponds to a lookup table for The method of embodiment 18. Aspect 20 Periodicity T>T max For the model parameters, ·Residual periodicity T r [T max -1,T max ], we subtract an integer value from T to find the residual periodicity T r and the step of determining; The lookup table for determining the prediction coefficients is calculated based on the residual periodicity T r and determining the value as a lookup table for 20. The method of embodiment 19. Aspect 21 Periodicity T <T min For the model parameters, The lookup table for determining the prediction coefficients is min selecting as a lookup table for a lookup parameter for identifying a selected lookup table entry that provides said prediction coefficient, the ratio T min and a scaling step using / T; determining the prediction coefficients using a selected lookup table and scaled lookup parameters; 21. The method of any one of embodiments 19 to 20. Aspect 22 determining a prediction mask (203, 205) indicative of a plurality of previous samples in a plurality of prediction mask support subbands, the plurality of prediction mask support subbands including at least one of the plurality of subbands different from the first subband; determining prediction coefficients to be applied to the previous samples based on the signal model, based on the model parameters, and based on the analysis filter bank; determining an estimate of the first sample by applying the plurality of prediction coefficients to the plurality of previous samples, respectively; 22. The method of any one of embodiments 1 to 21. Aspect 23 23. The method of embodiment 22, wherein determining the estimate of the first sample comprises determining a sum of the plurality of previous samples weighted by the plurality of respective prediction coefficients. Aspect 24 The plurality of subbands have equal subband spacing; the first subband is one of the plurality of subbands; 24. The method of any one of embodiments 1 to 23. Aspect 25 the analysis filters of the analysis filter bank are shift-invariant with respect to each other; and / or the analysis filters of the analysis filter bank include a common window function; and / or the analysis filters of the analysis filter bank comprise differently modulated versions of the common window function; and / or the common window function is modulated using a cosine function; and / or the common window function has a finite duration K; and / or the analysis filters of the analysis filter bank form an orthogonal basis; and / or the analysis filters of the analysis filter bank form an orthonormal basis; and / or the analysis filter bank comprises a cosine-modulated filter bank; and / or the analysis filter bank is a critically sampled filter bank; and / or the analysis filter bank includes a lapped transform; and / or the analysis filter bank includes one or more of an MDCT, QMF, ELT transform; and / or the analysis filter bank includes a modulation structure; 25. The method of any one of embodiments 1 to 24. Aspect 26 1. A method for estimating a first sample of a first subband signal in a first subband of an audio signal, wherein the first subband signal of the audio signal is determined using an analysis filter bank having a plurality of analysis filters providing a plurality of subband signals in a plurality of subbands from the audio signal, the method comprising: determining a prediction mask (203, 205) indicative of a plurality of previous samples in a plurality of prediction mask support subbands, the plurality of prediction mask support subbands including at least one of the plurality of subbands different from the first subband; determining a plurality of prediction coefficients to be applied to the plurality of previous samples; determining an estimate of the first sample by applying the plurality of prediction coefficients to the plurality of previous samples, respectively; method. Aspect 27 the plurality of prediction mask support subbands are including the first subband; and / or including one or more of the plurality of subbands immediately adjacent to the first subband, 27. The method of embodiment 26. Aspect 28 the method further comprising determining model parameters of the signal model; determining the prediction coefficients includes determining the prediction coefficients based on the signal model, based on the model parameters, and based on the analysis filter bank; 28. The method of embodiment 26 or 27. Aspect 29 determining the plurality of prediction coefficients includes determining the plurality of prediction coefficients using a look-up table or an analytical function; the lookup table or the analytical function provides the prediction coefficients as a function of a parameter derived from the model parameters; the lookup table or the analysis function is predetermined based on the signal model and based on the analysis filter bank; 29. The method of embodiment 28. Aspect 30 A method for encoding an audio signal, comprising: determining a plurality of subband signals from the audio signal using an analysis filter bank having a plurality of analysis filters; estimating samples of the plurality of subband signals using the method of any one of aspects 1-29, thereby providing a plurality of estimated subband signals; determining samples of a plurality of prediction error subband signals based on corresponding samples of the plurality of subband signals and samples of the plurality of estimated subband signals; quantizing the plurality of prediction error subband signals; indicating the plurality of quantized prediction error subband signals and indicating one or more parameters used to estimate the plurality of estimated subband signal samples; generating an encoded audio signal; method. Aspect 31 1. A method of decoding an encoded audio signal, the encoded audio signal indicating a plurality of quantized prediction error subband signals and one or more parameters to be used to estimate samples of a plurality of estimated subband signals, the method comprising: dequantizing the plurality of quantized prediction error subband signals to provide a plurality of dequantized prediction error subband signals; estimating samples of the plurality of estimated subband signals using the method of any one of aspects 1 to 29; determining samples of a plurality of decoded subband signals based on corresponding samples of the plurality of estimated subband signals and samples of the plurality of dequantized prediction error subband signals; determining a decoded audio signal from the plurality of decoded subband signals using a synthesis filter bank including a plurality of synthesis filters; method. Aspect 32 1. A system (103, 105) configured to estimate one or more first samples of a first subband signal of an audio signal, wherein the first subband signal of the audio signal is determined using an analysis filter bank having a plurality of analysis filters providing a plurality of subband signals from the audio signal, the system comprising: a predictor calculator configured to determine model parameters of a signal model and to determine one or more prediction coefficients to be applied to one or more previous samples of a first decoded subband signal derived from the first subband signal, the one or more prediction coefficients being determined based on the signal model, based on the model parameters, and based on the analysis filter bank, and time slots of the one or more previous samples being earlier than time slots of the one or more first samples; a subband predictor configured to determine estimates of the one or more first samples by applying the one or more prediction coefficients to the one or more previous samples, system. Aspect 33 1. A system (103, 105) configured to estimate one or more first samples of a first subband signal of an audio signal, the first subband signal corresponding to a first subband, the first subband signal being determined using an analysis filter bank having a plurality of analysis filters each providing a plurality of subband signals within the plurality of subbands, the system comprising: a predictor calculator configured to determine a prediction mask (203, 205) indicative of a plurality of previous samples in a plurality of prediction mask support subbands, the plurality of prediction mask support subbands including at least one of the plurality of subbands different from the first subband, the predictor calculator further configured to determine a plurality of prediction coefficients to be applied to the plurality of previous samples; a subband predictor configured to determine estimates of the one or more first samples by applying the plurality of prediction coefficients to the plurality of previous samples, respectively; system. Aspect 34 1. An audio encoder configured to encode an audio signal, comprising: an analysis filter bank configured to determine a plurality of subband signals from the audio signal using a plurality of analysis filters; 34. The system of any one of aspects 32-33, configured to estimate samples of the plurality of subband signals, thereby providing a plurality of estimated subband signals; a differencing unit configured to determine samples of a plurality of prediction error subband signals based on corresponding samples of the plurality of subband signals and the plurality of estimated subband signals; a quantization unit configured to quantize the plurality of prediction error subband signals; a bitstream generation unit configured to generate an encoded audio signal indicative of the plurality of quantized prediction error subband signals and one or more parameters used to estimate samples of the plurality of estimated subband signals, Audio encoder. Aspect 35 1. An audio decoder configured to decode an encoded audio signal, the encoded audio signal indicating the plurality of quantized prediction error subband signals and one or more parameters used to estimate samples of a plurality of estimated subband signals, the audio decoder comprising: an inverse quantizer configured to dequantize the plurality of quantized prediction error subband signals, thereby providing a plurality of dequantized prediction error subband signals; 34. The system of aspect 32 or 33, configured to estimate samples of the plurality of estimated subband signals; a summation unit configured to determine samples of a plurality of decoded subband signals based on corresponding samples of the plurality of estimated subband signals and based on samples of the plurality of dequantized prediction error subband signals; a synthesis filter bank configured to determine a decoded audio signal from the plurality of decoded subband signals using a plurality of synthesis filters; Audio decoder.
Claims
1. 1. A method performed by an audio signal processing apparatus for determining a current sample of a subband signal, the subband signal corresponding to one of a plurality of subbands of a subband domain representation of an audio signal, the method comprising: receiving preliminary samples of the subband signals; determining signal model data including model parameters; determining prediction coefficients to be applied to previous samples of the subband signal, the prediction coefficients being determined using a look-up table and / or an analytical function in response to the model parameters; determining estimated samples of the subband signals, wherein determining the estimated samples comprises applying the prediction coefficients to the previous samples; combining the estimated sample of the subband signal with the preliminary sample of the subband signal to determine a current sample of the subband signal; Including; the plurality of subbands have equal subband spacing; method.
2. 1. An audio signal processing apparatus configured to determine a current sample of a subband signal, the subband signal corresponding to one of a plurality of subbands of a subband domain representation of an audio signal, the audio signal processing apparatus comprising: receiving preliminary samples of the subband signals; determining signal model data including model parameters; determining prediction coefficients to be applied to previous samples of the subband signal, the prediction coefficients being determined using a look-up table and / or an analytical function in response to the model parameters; determining estimated samples of the subband signals, wherein determining estimated samples of the subband signals comprises applying the prediction coefficients to the previous samples; combining the estimated sample of the subband signal with the preliminary sample of the subband signal to determine a current sample of the subband signal; is configured to run the plurality of subbands have equal subband spacing; Audio signal processing device.
3. A non-transitory computer readable storage medium having a sequence of instructions that, when executed by a computer, causes the computer to perform the method of claim 1.