Decoder system, decoding method, and computer program
Complex predictive stereo coding in the frequency domain addresses the inefficiencies of current USAC systems at high bitrates by transforming stereo signals into efficient frequency domain representations, resulting in reduced computational complexity and delays for high-quality audio coding.
Patent Information
- Application Number
- JP2025039836
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2010-04-09
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-05
- Estimated Expiration
- 2031-04-06
AI Technical Summary
Current MPEG Unified Speech and Audio Coding (USAC) systems are computationally intensive and introduce delays at high bitrates, making them inefficient for high-quality stereo audio coding.
The proposed method employs complex predictive stereo coding using a coder and decoder system that transforms stereo signals into frequency domain representations of downmix and residual signals, allowing for efficient encoding and decoding without the need for additional QMF transforms.
This approach reduces computational complexity and delays, enabling high-quality stereo audio coding at high bitrates while preserving the advantages of unified stereo coding in MPEG USAC systems.
Smart Images

Figure 2025085702000001_ABST
Abstract
Description
[Technical field]
[0001] The invention disclosed herein relates generally to stereo audio coding, and more particularly to a technique for stereo coding using complex prediction in the frequency domain. [Background technology]
[0002] Joint coding of the left (L) and right (R) channels of a stereo signal allows for more efficient coding compared to coding L and R independently. A common approach to joint stereo coding is mid / side (M / S) coding, where the mid (M) signal is constructed by adding the L and R signals, e.g. the M signal is
number
number
[0003] The Moving Picture Experts Group (MPEG) Advanced Audio Coding (AAC) standard (see standard document ISO / IEC 13818-7) allows for a time- and frequency-variable choice between L / R and M / S stereo coding. Thus, a stereo encoder can apply L / R coding to some frequency bands of a stereo signal, while M / S coding is used to encode other frequency bands of the stereo signal (frequency-variable). Furthermore, the encoder can switch between L / R and M / S coding in time (time-variable). In MPEG AAC, the stereo encoding is performed in the frequency domain, more specifically in the Modified Discrete Cosine Transform (MDCT) domain. This allows for an adaptive choice between L / R or M / S coding in a frequency- and time-variable manner.
[0004] Parametric stereo coding is a method to efficiently code a stereo audio signal as a mono signal and a small amount of side information, which are the stereo parameters. It is part of the MPEG-4 Audio standard (see standard document ISO / IEC 14496-3). The mono signal can be encoded using any audio coder. The stereo parameters are embedded in the ancillary part of the mono bit stream, making it fully forward and backward compatible. At the decoder, the mono signal is first decoded and then the stereo signal is reconstructed using the stereo parameters. A decorrelated version of the decoded mono signal has zero cross-correlation with the mono signal. This decorrelating signal is generated by a decorrelator, for example by a suitable all-pass filter including a delay line. Essentially, the decorrelated signal has the same spectral and temporal energy distribution as the mono signal. The mono signal together with the decorrelated signal is input to an upmix process. This process is controlled by the stereo parameters and reconstructs the stereo signal. For more information, see [1].
[0005] MPEG Surround (MPS; see ISO / IEC 23003-1 and non-patent document 2) combines the principles of parametric stereo coding with those of residual coding, replacing a decorrelated signal with a transmitted residual, improving the perceived sound quality. Residual coding is performed by downmixing a multichannel signal and, optionally, extracting spatial cues. In the downmix process, a residual signal representative of the error signal is calculated, encoded and transmitted. The residual signal replaces the decorrelated signal at the decoder. In the hybrid approach, the residual signal replaces the decorrelated signal in certain frequency bands, preferably in the relatively low bands.
[0006] In current MPEG Unified Speech and Audio Coding (USAC) systems, two examples of which are shown in Figure 1, the decoder has a complex-valued Quadrature Mirror Filter (QMF) bank located downstream of the core decoder. The QMF representation obtained at the output of this filter bank is complex-valued and therefore oversampled by a factor of two, and can be configured as a downmix signal (i.e. mid signal) M and a residual signal D. This can be achieved by using an upmix matrix with complex-valued components. The L and R signals (in the QMF domain) are
number
[0007] The above coding scheme is well suited for low bitrates, typically below 80 kb / s, but is not optimal for high bitrates in terms of computational complexity. More specifically, at high bitrates, SBR tools are generally not used (as they do not improve coding efficiency). Secondly, decoders without an SBR stage use QMF filterbanks, which are computationally intensive and introduce delays due to the presence of complex-valued upmix matrices (for a frame length of 1024 samples, the QMF analysis / synthesis filterbank introduces a delay of 961 samples). This clearly indicates the need for a more efficient coding scheme. [Prior art documents] [Non-patent literature]
[0008] [Non-Patent Document 1] "Low Complexity Parametric Stereo Coding in MPEG-4" by H. Purnhagen, Proc. of the 7th Int. Conference on Digital Audio Effects (DAFx'04), Naples, Italy, October 5-8, 2004, pages 163-168 [Non-Patent Document 2] "MPEG Surround - The ISO / MPEG Standard for Efficient and Compatible Multi-Channel Audio Coding" by J. Herre et al., Audio Engineering Convention Paper 7084, 122 <nd>Convention, May 5-8, 2007 Summary of the Invention [Means for solving the problem]
[0009] It is an object of the present invention to provide a method and device for stereo coding that is computationally efficient even in the high bitrate range.
[0010] The invention achieves this object by providing a coder and decoder, a coding and decoding method and a computer program product for encoding and decoding, respectively, as defined in the independent claims. The dependent claims define embodiments of the invention.
[0011] In a first aspect, the present invention provides a decoder system for providing a stereo signal by complex predictive stereo coding, comprising: an upmix stage configured to generate the stereo signal based on first frequency domain representations of a downmix signal (M) and a residual signal (D), each first frequency domain representation having first spectral components representative of a spectral content of a corresponding signal expressed in a first subspace of a multidimensional space, the upmix stage comprising: a module for calculating a second frequency domain representation of the downmix signal based on a first frequency domain representation of the downmix signal, the second frequency domain representation having second spectral components representative of a spectral content of the signal expressed in a second subspace of the multidimensional space, the second subspace including a part of the multidimensional space that is not included in the first subspace; an upmix stage having a weighted adder for calculating a side signal (S) based on first and second frequency domain representations of the downmix signal encoded in the bitstream signal, a first frequency domain representation of the residual signal, and complex prediction coefficients (α); a sum / difference stage for calculating the stereo signal based on a first frequency domain representation of the downmix signal and the side signal, The upmix stage may furthermore be operable in a pass-through mode, in which the downmix signal and the residual signal are fed directly to the sum-and-difference stage.
[0012] In a second aspect, the present invention provides a system comprising: 1. An encoder system for encoding a stereo signal from a bitstream signal by complex predictive stereo coding, comprising: an estimator for estimating complex prediction coefficients; (a) an encoding stage operable to transform said stereo signal into a frequency domain representation of a downmix signal and a residual signal having a relationship determined by values of said complex prediction coefficients; A multiplexer receives the output from the encoding stage and the estimator and encodes it into the bitstream signal.
[0013] In a third and fourth aspect of the present invention, there are provided methods for encoding a stereo signal into a bitstream and for decoding the bitstream into at least one stereo signal. The technical features of each method are similar to the technical features of the encoder system and the decoder system, respectively. In a fifth and sixth aspect, the present invention provides a computer program product comprising instructions for carrying out each method on a computer.
[0014] The present invention benefits from the advantages of unified stereo coding in the MPEG USAC system. These advantages are preserved even at high bit rates without a significant increase in the computational complexity associated with QMF-based approaches, where SBR is not typically used, and this is possible because the critically sampled MDCT transform, which is the basis of the MPEG USAC transform, can also be used in complex predictive stereo coding according to the present invention, at least in cases where the code audio bandwidth of the downmix and residual channels is the same and the upmix process does not involve decorrelation. This means that an additional QMF transform is no longer necessary. Representative embodiments of complex predictive stereo coding in the QMF domain significantly increase the number of operations per unit time compared to conventional L / R or M / S stereo. The coding device according to the present invention therefore appears to be competitive at such bit rates, since it provides high sound quality with a modest computational load.
[0015] As those skilled in the art will be aware, the fact that the upmix stage can also operate in pass-through mode allows the decoder to adaptively decode by conventional direct or joint encoding and complex predictive encoding, depending on the decision of the encoder side. Thus, the decoder can ensure that the quality level is at least kept at the same level, if not more aggressively than conventional direct L / R stereo encoding or joint M / S stereo encoding. Thus, the decoder according to this aspect of the invention can be considered as a superset of the background art from a functional point of view.
[0016] The advantage over QMF-based predictive coded stereo is that it allows perfect reconstruction of the signal (except for the quantization error, which can be arbitrarily small).
[0017] Thus, the present invention provides a coding device for transform-based stereo coding with complex prediction. Preferably, the device according to the present invention is not limited to complex prediction stereo coding, but can also operate with direct L / R stereo coding according to the background art or simultaneous M / S stereo coding, allowing the most suitable coding method to be selected for a specific application or at a specific time.
[0018] An oversampled representation (e.g. a complex representation) of the signal, including both the first and second spectral components, is used as the basis for the complex prediction according to the invention, and thus a module for computing such an oversampled representation is configured in the encoder and decoder systems according to the invention. The spectral components refer to a first and a second subspace of a multidimensional space. It is a set of time-dependent functions of a given time length (e.g. the length of a predefined time frame) sampled with a finite sampling frequency. It is well known that functions in this multidimensional space can be approximated by a finite weighted sum of basis functions.
[0019] As will be clear to those skilled in the art, an encoder arranged to cooperate with a decoder is equipped with equivalent modules providing an oversampled representation on which the predictive coding is based, so as to allow a faithful reproduction of the encoded signal. Such equivalent modules are identical or similar modules or modules having the same or similar transfer characteristics. In particular, the encoder and decoder modules may each be similar or dissimilar units executing computer programs performing equivalent mathematical operations.
[0020] In some embodiments of the decoder system or encoder system, the first spectral components have real values represented in a first subspace and the second spectral components have imaginary values represented in a second subspace. Together the first and second spectral components comprise a complex spectral representation of the signal. The first subspace is a linear span of a first set of basis functions and the second subspace is a linear span of a second set of basis functions, some of which are linearly independent of the first set of basis functions.
[0021] In one embodiment, the module for computing the complex representation is a real-to-imaginary transform, i.e. a module for computing the imaginary value of the spectrum of the discrete-time signal based on a real spectral representation of the signal, said transform being based on exact or approximate mathematical relations, such as formulas from harmonic analysis or heuristic relations.
[0022] In an embodiment of the decoder system or the encoder system, the first spectral component is determined by a time-to-frequency domain transformation of the discrete time domain signal, preferably by a Fourier transform, such as a discrete cosine transform (DCT), a modified discrete cosine transform (MDCT), a discrete sine transform (DST), a modified discrete sine transform (MDCT), a fast Fourier transform (FFT), a prime-factor-based Fourier algorithm, etc. In the first four cases, the second spectral component is determined by a DST, an MDST, a DCT, and an MDCT, respectively. As is well known, the linear span of a cosine periodic over a unit period constitutes a subspace that is not completely contained in the linear span of a sine periodic over the same period. Preferably, the first spectral component is determined by an MDCT and the second spectral component is determined by an MDST.
[0023] In one embodiment, the decoder system includes at least one temporal noise shaping module (TNS module, i.e. TNS filter), which is arranged upstream of the upmix stage. Generally speaking, the use of TNS improves the perceived sound quality of signals with transient-like components, which also applies to the embodiment of the inventive decoder system with TNS. In conventional L / R and M / S stereo coding, the TNS filter is applied as the last processing step in the frequency domain, just before the inverse transform. However, in the case of complex predictive stereo coding, it is often advantageous to apply the TNS filter to the downmix signal and the residual signal, i.e. before the upmix matrix. In other words, TNS is applied to a linear combination of the left and right channels, which has several advantages. First, it turns out that in certain situations TNS is advantageous, for example, only for the downmix signal. Secondly, for the residual signal TNS filtering can be omitted, which means an economical use of the available bandwidth. The TNS filter coefficients only need to be transmitted for the downmix signal. Secondly, the computation of an oversampled representation of the downmix signal (e.g. MDST data are derived from MDCT data to construct a complex frequency domain representation), which is necessary in complex predictive coding, requires that a time domain representation of the downmix signal is computable. This means that the downmix signal is available as a time sequence of MDCT spectra, preferably derived uniformly. If a TNS filter is applied in the decoder after the upmix matrix that transforms the downmix / residual representation into a left / right representation, only a sequence of TNS residual MDCT spectra of the downmix signal is obtained. This makes the efficient computation of the corresponding MDST spectrum very difficult, especially when the left / right channels use TNS filters with different characteristics.
[0024] It is emphasized that the availability of a time sequence of MDCT spectra is not an absolute criterion for obtaining a MDST representation that is fit to serve as a basis for complex predictive coding. In addition to experimental evidence, this fact can generally be explained by the fact that TNS is only applied at high frequencies, e.g. higher than a few kilohertz, so that the residual signal filtered by TNS approximately corresponds to the low-frequency unfiltered residual signal. Thus, the invention can be implemented as a decoder with complex predictive stereo coding, in which the TNS filter is placed other than upstream of the upmix stage, as will be explained below.
[0025] In one embodiment, the decoder system includes at least one separate TNS module arranged downstream of the upmix stage. A selector device allows the selection of a TNS module upstream of the upmix stage or a TNS module downstream of the upmix stage. Under certain circumstances, the calculation of a complex frequency domain representation does not require that a time domain representation of the downmix signal can be calculated. Furthermore, as mentioned above, the decoder can selectively operate in a direct or joint coding mode without applying complex predictive coding, making it more suitable to use a TNS module in its traditional place, i.e. as one of the last processing steps in the frequency domain.
[0026] In one embodiment, the decoder system is configured to save processing resources and possibly energy by deactivating a module that computes a second frequency domain representation of the downmix signal. The downmix signal is partitioned into consecutive time blocks, each time block being associated with a value of a complex prediction coefficient. This value is determined for each time block by an encoder cooperating with the decoder. Furthermore, in this embodiment, the module that computes the second frequency domain representation of the downmix signal is configured to deactivate itself if, for a given time block, the absolute value of the imaginary part of the complex prediction coefficient is zero or is smaller than a predefined tolerance value. Deactivation of the module means that it does not compute a second frequency domain representation of the downmix signal for this time block. In the absence of deactivation, the second frequency domain representation (e.g. a set of MDST coefficients) is multiplied by zero or a number approximately of the same order as the decoder's machine epsilon (rounded off) or other suitable threshold.
[0027] In a further development of said embodiment, the saving of processing resources is performed at sub-levels of the time block to which the downmix signal is partitioned. For example, such sub-levels within the time block are frequency bands, and the encoder determines values of complex prediction coefficients for each frequency band within the time block. Similarly, the method for generating a second frequency domain representation is configured to suppress operations for frequency bands within the time block in which the complex prediction coefficient is zero or has a magnitude smaller than a tolerance value.
[0028] In one embodiment, the first spectral components are transform coefficients arranged in time blocks of transform coefficients, each block being generated by application of a transform to a time segment of the time domain signal. Furthermore, the module for computing a second frequency domain representation of the downmix signal comprises: determining a first intermediate component from the first spectral component; forming a combination of the first spectral components according to at least a portion of an impulse response to obtain a second intermediate component; configured to determine a second spectral component from the second intermediate component. This procedure allows the calculation of the second frequency domain representation directly from the first frequency domain representation, as described in detail in US Patent No. 6,980,933 B2, in particular columns 8-28, and in particular equation 41. As those skilled in the art will be aware, the calculation is not performed in the time domain, as opposed to, for example, an inverse transformation followed by a different transformation.
[0029] It is estimated that for the embodiment of the complex predictive stereo coding according to the invention, the computational complexity increases only slightly compared to conventional L / R or M / S stereo (significantly less than the increase caused by complex predictive stereo coding in the QMF domain). This type of embodiment, which includes an exact calculation of the second spectral components, introduces a delay that is only a few percent longer than that introduced by the QMF-based embodiment (assuming a time block length of 1024 samples, compared to a delay of 961 samples of the QMF analysis / synthesis filter bank).
[0030] Advantageously, at least in some of the above embodiments, the impulse response is adapted to a transformation required for the first frequency domain representation, or more precisely, for its frequency response characteristics.
[0031] In some embodiments, the first frequency domain representation of the downmix signal is obtained by a transformation applied to one or more analysis window functions (or cut-off functions, e.g. rectangular window, sine window, Kaiser-Bessel window, etc.), the purpose of which is to achieve a temporal segmentation without introducing dangerous noise loudness or undesirable changes to the spectrum. Possibly, such window functions are partially overlapping. Then, preferably, the frequency response characteristics of the transformation depend on the characteristics of said one or more analysis window functions.
[0032] With further reference to the embodiment characterized by the calculation of the second frequency domain representation in the frequency domain, the computation load can be reduced by using an approximate second frequency domain representation. Such an approximation can be achieved by not requiring perfection in the information on which the calculation is based. For example, according to the teaching of US Pat. No. 6,980,933 B2, first frequency domain data from three time blocks, namely a block contemporaneous with the output block, a preceding block and a following block, are required for the exact calculation of the second frequency domain representation of the downmix signal in one block. For the purpose of complex predictive coding according to the invention, a good approximation can be obtained by omitting or replacing with zero data from the following block and / or the preceding block (caused by the operation of the module, i.e. not contributing to the delay), so that the calculation of the second frequency domain representation is based on only one or two time blocks. It should be noted that the omission of input data implies a rescaling of the second frequency domain representation, for example in the sense that it no longer represents the same power, but as stated above it can still be used as the basis for complex predictive coding, as long as it is calculated in an equivalent way on both the encoder and decoder side. Indeed, this kind of rescaling is compensated for by a corresponding change in the values of the prediction coefficients.
[0033] Yet another approximation method for calculating the spectral components constituting part of the second frequency domain representation of the downmix signal includes combining at least two components from the first frequency domain representation, the latter components being adjacent in time and / or frequency. Alternatively, the latter components can be combined by finite impulse response (FIR) filtering in a relatively small number of steps. For example, in a system using a time block size of 1024, such an FIR filter may include 2, 3, 4 etc. taps. A description of this type of approximate calculation method can be found, for example, in US Patent Application Publication No. 2005 / 0197831 A1. If a window function is used that gives a relatively small weight to the vicinity of each time block boundary, for example a non-rectangular function, it is advantageous to base the second spectral components of a time block only on a combination of the first spectral components of the same time block, but the same amount of information about the outermost components will not be available. The approximation errors that may result from such a practice can be limited to some extent or hidden by the shape of the window function.
[0034] In an embodiment of a decoder designed to output a time-domain stereo signal, there is the possibility to switch between direct or simultaneous stereo coding and complex predictive coding. This can be achieved by providing: · A switch that can selectively operate as a pass-through stage (which does not alter the signal) or as a sum-difference transformation; an inverse transform stage performing a frequency-time transform; and A selector device for inputting the directly (or simultaneously) coded signal or the signal coded by complex prediction to the inverse transform stage. As one skilled in the art will appreciate, with such flexibility on the part of the decoder, the encoder has the freedom to choose between conventional direct or joint coding and complex predictive coding. Thus, this embodiment can ensure that the quality level of the conventional direct L / R stereo coding or joint M / S stereo coding is at least maintained, if not exceeded. Thus, the decoder according to this embodiment can be considered as a superset of the related art.
[0035] Another group of embodiments of the decoder system performs the calculation of the second spectral components in a second frequency domain representation via the time domain. More precisely, it applies an inverse transform of the transform that obtained (or can obtain) the first spectral components, and then performs a different transform having the second spectral components as output. In particular, it performs an inverse MDCT followed by an MDST. In such embodiments, to reduce the number of transforms and inverse transforms, the output of the inverse MDCT is sent to the MDST and to an output terminal of the decoding system (possibly preceded by further processing steps).
[0036] It is estimated that for an embodiment of complex predictive stereo coding according to the invention, the computational complexity increases only slightly compared to conventional L / R or M / S stereo (significantly less than the increase caused by complex predictive stereo coding in the QMF domain).
[0037] As a further development of the embodiment mentioned in the previous paragraph, the upmix stage may comprise a further inverse transform stage processing the side signal, and supplying to the sum-and-difference stage a time-domain representation of the side signal generated by said further inverse transform stage and a time-domain representation of the downmix signal generated by said inverse transform. Again, advantageously from a computational complexity point of view, the latter signal is supplied both to the above-mentioned sum-and-difference stage and to the different transform stage.
[0038] In one embodiment, a decoder designed to output a time-domain stereo signal is capable of switching between direct L / R stereo coding or simultaneous M / S stereo coding, and complex predictive stereo coding. This can be achieved by comprising: Switches that can operate as pass-through stages or as sum-difference stages; A further inverse transform stage that computes the time domain representation of the side signal; a selector device for connecting the inverse transform stage to a further sum / difference stage connected to a point upstream of the upmix and downstream of the switch (preferably when the switch is activated and functions as a pass filter, such as when decoding a stereo signal generated by complex predictive coding) or to a combination of the downmix signal from the switch and the side signal from the weighted adder (preferably when the switch is activated and functions as a sum / difference stage, such as when decoding a directly coded stereo signal). As those skilled in the art will be aware, this gives the encoder the freedom to choose between conventional direct or simultaneous encoding and complex predictive encoding, i.e., to guarantee a sound quality level at least equal to that of direct or simultaneous stereo encoding.
[0039] In one embodiment, the encoder system according to the second aspect of the invention comprises an estimator for estimating complex prediction coefficients with the aim of reducing or minimizing the signal power or the average signal power of the residual signal. The minimization is performed over time, preferably over the time segment or time block or time frame to be coded. The squared amplitude can be a measure of the instantaneous signal power, and the integral of the squared amplitude over a time interval can be a measure of the average signal power in that time interval. Advantageously, the complex prediction coefficients can be determined for each time block and for each frequency band, i.e. their values are set to reduce the average power (i.e. total energy) of the residual signal in that time block and frequency band. In particular, the module for estimating parametric stereo coding parameters such as IID, ICC and IPD or similar parameters provides an output from which the complex prediction coefficients can be calculated according to mathematical relations known to those skilled in the art.
[0040] In one embodiment, the encoding stage of the encoder system further functions as a pass-through stage to enable direct stereo encoding. In situations where direct stereo encoding is expected to provide higher sound quality, by selecting this, the encoder system can ensure that the encoded stereo signal has at least the same sound quality as direct encoding. Similarly, in situations where the high computational load caused by complex predictive encoding is undesirable even if the sound quality is significantly improved, the encoder system has readily available options to save computational resources. The decision between joint, direct real predictive encoding and complex predictive encoding in a coder is generally based on the principle of rate / distortion optimization.
[0041] In one embodiment, the encoder system comprises a module for calculating a second frequency domain representation based directly on the first spectral components (i.e. without applying an inverse transform to the time domain and without using the time domain data of the signal). With respect to the corresponding embodiment of the decoder system described above, this module has a similar configuration, i.e. with a similar but different order of processing operations, and the encoder is configured to output data suitable for the input on the decoder side. For the purposes of describing this embodiment, it is assumed that the stereo signal to be encoded has a mid and a side channel or is transformed into this configuration, and the encoding stage is configured to receive a first frequency domain representation. The encoding stage comprises a module for calculating a second frequency domain representation of the mid channel. (The first and second frequency domain representations referred to here are as defined above; in particular, the first frequency domain representation may be an MDCT representation and the second frequency domain representation may be an MDST representation.) The encoding stage further comprises a weighting adder for calculating a residual signal, composed of the side signal and the two frequency domain representations of the mid signal, as a linear combination weighted by the real and imaginary parts of the complex prediction coefficients. The mid signal, or preferably its first frequency domain representation, is directly used as downmix signal. In this embodiment, furthermore, an estimator determines values of complex prediction coefficients with the aim of minimizing the power of the residual signal or the average signal power. The final operation (optimization) is performed by feedback control, where the estimator further receives, if necessary, the residual signal resulting from the current prediction coefficient values to be adjusted, or in a feedforward manner, by calculations performed directly on the left / right channels or on the mid / side channels of the original stereo signal. A feedforward method is preferred, where the complex prediction coefficients are calculated directly (in particular non-iteratively or non-feedback) based on the first and second frequency domain representation of the mid signal and the first frequency domain representation of the side signal. It should be noted that after the determination of the complex prediction coefficients, a decision is made to directly perform joint real predictive coding or complex predictive coding, taking into account the quality (preferably perceptual quality taking into account, for example, the signal-to-mask effect) obtained with each option.Therefore, the above statement should not be interpreted as meaning that there is no feedback mechanism in the encoder.
[0042] In one embodiment, the encoder system comprises a module for calculating a second frequency domain representation of the mid (or downmix) signal via the time domain. The implementation details for this embodiment are similar, at least as far as the calculation of the second frequency domain representation is concerned, and can be performed in the same way as for the corresponding decoder embodiment. In this embodiment, the encoding stage comprises: A sum / difference stage to convert the stereo signal into a mid channel and a side channel; a transformation step providing a frequency domain representation of the side channel and a complex-valued (i.e. oversampled) frequency domain representation of the mid channel; and · A weighted adder that calculates the residual signal using the complex prediction coefficients as weights. Here, the estimator receives the residual signal and determines complex prediction coefficients that reduce or minimize the power or average power of the residual signal, possibly in the form of a feedback control. However, preferably, the estimator receives the stereo signal to be coded and determines the prediction coefficients on the basis of that. Using a critically sampled frequency domain representation of the side channel is advantageous in terms of computational economy, since in this embodiment the side channel is not multiplied with complex numbers. Advantageously, the transformation stage may comprise an MDCT stage and an MDST stage arranged in parallel, both of which have as input a time domain representation of the mid channel. In this way, an oversampled frequency domain representation of the mid channel and a critically sampled frequency domain representation of the side channel are generated.
[0043] It should be noted that the methods and apparatus disclosed in this section may be adapted, with appropriate modifications within the ability of one of ordinary skill in the art, including routine experimentation, to encode signals having more than two channels. Modifications for such multi-channel operability may be made, for example, in accordance with Sections 4 and 5 of the article by J. Herre et al. cited above.
[0044] In further embodiments, features of two or more of the above embodiments may be combined, as long as they are not obviously complementary. The fact that two features are recited in different claims does not mean that they cannot be combined. Similarly, further embodiments may omit features that are not necessary or essential for the desired purpose. As an example, the decoding system according to the invention may be implemented without a dequantization stage if the encoded signal to be processed is not quantized or is already in a form suitable for processing in the upmix stage. [Brief description of the drawings]
[0045] The invention will be further illustrated by the embodiments described in the following sections with reference to the attached drawings. [Figure 1A] FIG. 1 is a block diagram showing a QMF-based decoder according to the background art. [Figure 1B] FIG. 1 is a block diagram showing a QMF-based decoder according to the background art. [Diagram 2] 1 is a block diagram showing an MDCT-based stereo decoder system with complex prediction according to one embodiment of the present invention, where the complex representations of the channels of the signal to be decoded are calculated in the frequency domain. [Diagram 3] 1 is a block diagram showing an MDCT-based stereo decoder system with complex prediction according to one embodiment of the present invention, in which the complex representations of the channels of the signal to be decoded are calculated in the time domain. [Figure 4] Figure 3 shows an alternative embodiment of the decoder system of Figure 2. The position of the active TNS stage is selectable. [Diagram 5] FIG. 13 is a block diagram showing an MDCT-based stereo encoder system with complex prediction according to an embodiment of another aspect of the present invention. [Figure 6] 1 is a block diagram showing an MDCT-based stereo encoder system with complex prediction according to one embodiment of the present invention, in which complex representations of the channels of a signal to be encoded are computed based on their time-domain representation. [Figure 7] Fig. 7 shows an alternative embodiment of the encoder system shown in Fig. 6, which is also operable in a direct L / R encoding mode. [Figure 8] FIG. 1 is a block diagram showing an MDCT-based stereo encoder system with complex prediction according to an embodiment of the present invention, in which complex representations of the channels of a signal to be encoded are calculated based on a first frequency domain representation thereof, and the system is also operable in a direct L / R encoding mode. [Figure 9] FIG. 8 illustrates an alternative embodiment of the encoder system shown in FIG. 7, further comprising a TNS stage disposed downstream of the encoding stage. [Figure 10] FIG. 9 illustrates an alternative embodiment of the portion labeled A in FIGS. 2 and 8. [Figure 11] Fig. 9 shows an alternative embodiment of the encoder system shown in Fig. 8. The system further comprises frequency domain modification devices respectively arranged upstream and downstream of the encoding stage. [Figure 12] Graph showing listening test results at 96 kb / s from 6 subjects, showing different complexity vs. sound quality trade-off options for computing or approximating the MDST spectrum. Here, data points labeled "+" indicate hidden references. "X" indicates a 3.5 kHz band-limited anchor. "*" indicates conventional stereo (M / S or L / R) with USAC. "□" indicates MDCT-domain unified stereo coding with complex prediction, with the imaginary part of the prediction coefficients disabled (i.e., with real-valued prediction that does not require MDST). "■" indicates MDCT-domain unified stereo coding with complex prediction, where an approximation of the MDST is computed using the current MDCT frame. "○" indicates MDCT-domain unified stereo coding with complex prediction, where an approximation of the MDST is computed using the current and previous MDCT frames. "●" indicates MDCT-domain unified stereo coding with complex prediction, where an approximation of the MDST is computed using the current, previous and next MDCT frames. [Figure 13] FIG. 13 shows the data of FIG. 12 as difference scores for MDCT domain unified stereo coding with complex prediction, which uses the current MDCT frame to compute an approximation of the MDCT. [Figure 14A] FIG. 2 is a block diagram illustrating an embodiment of a decoder system according to an embodiment of the present invention. [Figure 14B] FIG. 4 is a block diagram illustrating another embodiment of a decoder system according to an embodiment of the present invention. [Figure 14C] FIG. 11 is a block diagram illustrating yet another embodiment of a decoder system according to an embodiment of the present invention. [Figure 15] 4 is a flowchart illustrating a decoding method according to an embodiment of the present invention. [Figure 16] 4 is a flowchart illustrating an encoding method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0046] I. Decoder System FIG. 2 shows, in schematic block diagram form, a method for detecting at least one complex prediction coefficient value α=α R +iα I The MDCT representation of the stereo signal has M downmix channels and D residual channels. R and the imaginary part α I are quantized and / or coded jointly. However, preferably, the real and imaginary parts are quantized independently and uniformly, typically with a step size of 0.1 (a dimensionless number). According to the MPEG standard, the resolution of the frequency band used for the complex prediction coefficients does not have to be the same as the resolution of the scale factor band (sfb, i.e. a group of MDCT lines using the same MDCT quantization step size and quantization range). In particular, the frequency band resolution of the prediction coefficients is psychoacoustically relevant, such as the Bark scale. The demultiplexer 201 is configured to extract these MDCT representations and the prediction coefficients (part of the control information shown) from the bitstream provided to it. In fact, the bitstream contains encoded control information beyond the complex prediction coefficients, such as an instruction on whether to decode it in predictive or non-predictive mode, and TNS information, which includes the values of the TNS parameters used by the TNS (synthesis) filter of the decoder system. In case of using the same set of TNS parameters for multiple TNS filters, such as for both channels, it is more economical to receive this information in the form of bits indicating the identity of the set of parameters than to receive two sets of parameters separately. For example, information on whether the TNS should be applied before or after the upmix stage, based on a psychoacoustic evaluation of the two options, may also be included. Furthermore, the control information indicates the individual limited bandwidths of the downmix signal and the residual signal. For each channel, the frequency bands above the bandwidth limit are not decoded but are set to zero. In some cases, the energy content of the highest frequency bands is so small that they are already zero when quantized. In normal practice (see max_sfb parameter in the MPEG standard), the same bandwidth limit must be used for both the downmix signal and the residual signal. However, the residual signal has an energy content that is significantly more localized in the low frequency bands than the downmix signal. Therefore, by imposing a dedicated bandwidth limit on the residual signal, a reduction in the bit rate is possible without a significant loss of sound quality.For example, this is regulated by two independent max_sfb parameters encoded in the bitstream, one for the downmix signal and one for the residual signal.
[0047] In this embodiment, the MDCT representation of the stereo signal is segmented into consecutive time frames (i.e., time blocks) containing a fixed number of data points (e.g., 1024 points), one of a number of fixed numbers of data points (e.g., 128 points or 1024 points), or a variable number of points. As known to those skilled in the art, the MDCT is critically sampled. The output of the decoding system is a time-domain stereo signal having a left L channel and a right R channel, shown in the right part of the figure. The inverse quantization module 202 is configured to process the bitstream input to the decoding system and, if necessary, two bitstreams corresponding to the downmix channel and the residual channel obtained after demultiplexing the original bitstream. The inverse quantized channel signals are transformed by a transformation matrix
number
number
[0048] Assuming now that the switching assembly 203 is in pass-through mode, in this embodiment the dequantized channel signals are passed through the respective TNS filters 204. The TNS filters 204 are not essential for the operation of the decoding system and can be replaced by pass-through elements. After this, the signals are fed to a second switching assembly 205 having the same functionality as the upstream switching assembly 203. As described above, the output of the second switching assembly 205, when it receives an input signal and is set in pass-through mode, is the downmix channel signal and the residual channel signal. The downmix signal, represented by a time-continuous MDCT spectrum, is fed to a real-to-imaginary transform 206 configured to calculate the MDST spectrum of the downmix signal. In this embodiment, one MDST frame is based on three MDCT frames, one previous frame, one current (i.e. simultaneous) frame and one subsequent frame. It is symbolically understood that the input side of the real-to-imaginary transform 206 has a delay component (Z -1 ,Z) are shown.
[0049] The MDST representation of the downmix signal obtained from the real-to-imaginary transform 206 is the imaginary part of the prediction coefficients α I The real part of the prediction coefficient α R and is added to the MDCT representation of the downmix signal weighted by the MDCT representation of the residual signal. The two additions and multiplications are performed by adders and multipliers that (functionally) constitute weighted adders 210, 211. These are supplied with the values of the complex prediction coefficients α that were encoded in the bitstream originally received by the decoder system. The complex prediction coefficients are determined one for each time frame. They may be determined more frequently, one for each frequency band in the frame. The frequency bands are psychoacoustically motivated partitions. As will be explained later in the present coding system, the complex prediction coefficients may be determined less frequently. The real-to-imaginary transform 206 is synchronized with the weighted adders so that the current MDST frame of the downmix channel signal is combined with the respective simultaneous MDCT frames of the downmix channel signal and the residual channel signal. The sum of these three signals is the side signal S=Re{αM}+D. In this formula, M includes both the MDCT and MDST representations of the downmix signal, i.e. M=M MDCT -iM MDST D = DMDCT is real valued. In this way, a stereo signal with downmix channels and side channels is obtained, and the sum-difference transform 207 converts this stereo signal into
number
[0050] A possible implementation of the real-to-imaginary transform 206 is described in detail in the applicant's U.S. Pat. No. 6,980,933 B2, as mentioned above. The transform can be expressed as a finite impulse response filter, according to Equation 41 described in said document. For example, for an even number of points,
number
number
[0051] It is possible to further reduce the amount of input data on which the calculations are based. For illustrative purposes, the real-to-imaginary transform 206 and its upstream connections, indicated in the diagram by "A", may be replaced by simplified variants, two of which, A' and A'', are shown in Figure 10. Variation A' gives an approximation of the imaginary representation of the signal, where the MDST calculation only considers the current and previous frames. Referring to the formula above in this paragraph, we consider X for p=0,...,N-1. III This is done by setting p = 0 (index III indicates the later time frame). Variation A' does not require the MDCT spectrum of the later frame as input, so the MDST calculation does not introduce a time delay. Obviously, this approximation somewhat reduces the accuracy of the obtained MDST signal, but it also implies a reduction in the energy of this signal. As a nature of predictive coding, the latter is I This can be fully compensated for by increasing
[0052] Variation A'' is shown in Figure 10. It uses only the MDCT data of the current time frame as input. The MDST representation obtained by Variation A'' is less accurate than that obtained by Variation A'. On the other hand, Variation A'' works with zero delay like Variation A' and has low computational complexity. As mentioned before, as long as the same approximation is used in the encoder and decoder systems, the waveform coding properties are not affected.
[0053] It should be noted that, regardless of whether modification A, A′ or A″ or any of their further developments is used, the imaginary part of the complex prediction coefficients of the MDST spectrum is non-zero, i.e., α I Only the part where ≠ 0 needs to be calculated. In practical situations, this is the absolute value of the imaginary part of the coefficient |α I can be interpreted to mean that |α is greater than a predetermined threshold, which is related to the unit round-off of the hardware used. If the imaginary parts of the coefficients for all frequency bands in a time frame are zero, then there is no need to compute MDST data for that frame. Thus, again, real-to-imaginary transform 206 reduces |α by not producing an MDST output. I The second switching assembly 205 is configured to respond when the value of | is very small, which allows to save computational resources. However, in an embodiment where more than one frame is used to generate one frame of MDST data, the units upstream of the transform 206 must continue to operate even though the MDST spectrum is not needed, in particular the second switching assembly 205 must continue to forward the MDCT spectrum, so that there is enough input data for the real-to-imaginary transform 206 when the next time frame associated with a non-zero prediction coefficient occurs. This is of course the next time block.
[0054] Returning to FIG. 2, the functionality of the decoder system has been described assuming that both switching assemblies 203, 205 are set in pass-through mode. As described herein, the decoder system can also decode signals that are not predictively coded. For this application, the second switching assembly 205 is set in sum-and-difference mode and the selector device 208 is set in the lower position as shown in the figure, so that the signal is input directly to the inverse transform 209 from a source point between the TNS filter 204 and the second switching assembly 205. For correct decoding, the signal has an L / R format at the source point appropriately. Therefore, it is preferable to set the second switching assembly 205 in sum-and-difference mode when decoding non-predictively coded stereo signals in order to always provide the correct mid (i.e., downmix) signal to the real-to-imaginary transform (rather than, for example, more simply, the left signal). As mentioned above, predictive coding can be replaced by conventional direct coding or simultaneous coding of multiple frames, based, for example, on a data rate vs. sound quality decision. The result of such a decision is sent from the encoder to the decoder in various ways, for example by the value of a dedicated indicator bit in each frame or by the presence or absence of a prediction coefficient value. Having established these facts, the role of the first switching assembly 203 is easily realised. In fact, in the non-predictive coding mode, the decoder system can process both signals with direct (L / R) stereo coding and signals with simultaneous (M / S) coding. By operating the first switching assembly 203 either in pass-through mode or in sum-and-difference mode, it is possible to ensure that a source point is always provided together with the directly coded signal. Obviously, when the switching assembly 203 functions in the sum-and-difference stage, it converts the input signal in M / S format into an output signal in L / R format (supplied to the optional TNS filter 204).
[0055] The decoder system receives a signal indicating whether a time frame is to be decoded by the decoder system in predictive or non-predictive coding mode. The non-predictive mode is signaled by the value of a dedicated indicator bit in each frame or by the presence or absence (or value of zero) of prediction coefficients. The predictive mode can be signaled similarly. A particularly advantageous embodiment allows for an overhead-free fallback, but makes use of a reserved fourth value of the 2-bit field ms_mask_present (MPEG-2 AAC, see ISO / IEC 13818-7 document). This is transmitted every time frame and is specified as follows: [Table 1] By redefining the value 11 to mean "complex predictive coding", the decoder can operate in all legacy modes, in particular the M / S and L / R coding modes, without compromising bitrate, and can receive a signal indicating the complex predictive coding mode of the associated frame.
[0056] Fig. 4 shows a decoder system in a general configuration, similar to that shown in Fig. 2, but including at least two different configurations. Firstly, the system of Fig. 4 includes switches 404, 411 allowing the application of processing steps including frequency domain modifications upstream and / or downstream of the upmix stage. This is realized on the one hand by a first set of frequency domain modifiers 403 (depicted in this figure as TNS synthesis filters) provided with a first switch 404 downstream of the inverse quantization module 401 and the first switching assembly 402 and upstream of a second switching assembly 405 arranged immediately upstream of the upmix stages 406, 407, 408, 409. On the other hand, the decoder system includes a second set of frequency domain modifiers 410 provided with a second switch 411 downstream of the upmix stages 406, 407, 408, 409 and upstream of the inverse transform stage 412. Advantageously, as shown in the figure, each frequency domain modifier is arranged in parallel with a pass-through line connected upstream to the input of the frequency domain modifier and downstream to the associated switch. This configuration allows the frequency domain modifier to always be fed with signal data, allowing processing in the frequency domain based on more time frames than just the current one. The decision to apply either the first set of frequency domain modifiers 403 or the second set of frequency domain modifiers 410 may be made by the encoder (sent in the bitstream), or based on whether predictive coding is applied, or on other criteria suitable for the practical situation. As an example, if the frequency domain modifiers are TNS filters, the first set 403 is advantageous for use with certain types of signals, while the second set 410 is advantageous for use with other types of signals. If the result of this selection is encoded in the bitstream, the decoder system will activate each set of TNS filters accordingly.
[0057] To facilitate understanding of the decoder system shown in Fig. 4, it is explicitly noted that the decoding of a directly (L / R) encoded signal is performed when α=0 (implying that pseudo-L / R and L / R are the same and that the side and residual channels are not different), the first switching assembly 402 is in pass mode, the second switching assembly is in sum-and-difference mode and the signal is in M / S format between the second switching assembly 405 and the sum-and-difference stage 409 of the upmix stage. Since the upmix stage is then effectively a pass step, it does not matter whether the first set of frequency domain modifiers or the second set of frequency domain modifiers is activated (using the respective switches 404, 411).
[0058] FIG. 3 shows a decoder system according to an embodiment of the present invention, which represents a different approach to the provision of MDST data required for upmixing, in relation to the decoder systems of FIGS. 2 and 4. As with the decoder systems already described, the system of FIG. 3 comprises an inverse quantization module 301, a first switching assembly 302 operable in pass-through or sum-and-difference mode, and a TNS (synthesis) filter 303, all arranged in series from the input of the decoder system. The modules downstream of this point are selectively utilized by two second switches 305, 310. These second switches are preferably operated simultaneously, both in the up or down position as shown. At the output of the decoder system there is a sum-and-difference stage 312, immediately upstream of which there are two inverse MDCT modules 306, 311 which convert the MDCT domain representation of each channel into a time domain representation.
[0059] In complex predictive decoding, the decoder system is fed with a bitstream encoding the downmix / residual stereo signal and the complex prediction coefficients, the first switching assembly 302 is set in pass-through mode and the second switches 305, 310 are set in the up position. Downstream of the TNS filter, the two channels of the (dequantized, TNS filtered MDCT) stereo signal are processed differently. The downmix channel is fed on the one hand to a multiplier and adder 308, which multiplies the real part α of the prediction coefficients by R The MDCT representation of the downmix channel weighted by α is added to the MDCT representation of the residual channel on the other hand, which is fed to one of several MDCT transform modules 306. A time domain representation of the downmix channel M is output from the inverse MDCT transform module 306 and is fed to both the final sum-and-difference stage 312 and to the MDST transform module 307. This double use of the time domain representation of the downmix channel is advantageous in terms of computational complexity. The MDST representation of the downmix channel thus obtained is fed to a further multiplier and adder 309, which multiplies the imaginary part α of the prediction coefficients by I before adding this signal to the linear combination output from summer 308. Thus, the output of summer 309 is the side channel signal S=Re{αM}+D. Similarly, the multipliers and summers 308, 309 are coupled to the decoder system shown in Figure 2 to form a weighted multi-signal summer having as input the MDCT and MDST representations of the downmix signal, the MDCT representation of the residual signal and the complex prediction coefficient values. In this embodiment, downstream of this point, the only path remaining is through the inverse MDCT transform module 311, before the side channel signal is fed to the final sum and difference stage 312.
[0060] The required synchronicity in the decoder system can be achieved by applying the same transform length and window shape in both inverse MDCT transform modules 306, 311. This is already practiced in frequency selective M / S and L / R coding. Combining one embodiment of the inverse MDCT module 306 with one embodiment of the MDST module 307 will result in a delay of one frame. Therefore, five optional delay blocks 313 (or software instructions to this effect in case of computer implementation) are provided to allow the part of the system to the right of the dashed line to be delayed by one frame with respect to the part to the left, if necessary. Obviously, all intersections between the dashed line and the connecting lines are provided with delay blocks, except for the connection between the inverse MDCT module 306 and the MDST transform module 307, which introduces a delay that needs to be compensated.
[0061] The calculation of MDST data for one time frame requires data from one frame of the time domain representation. However, the inverse MDCT transform can be based on one frame (the current frame), two consecutive frames (preferably the previous and current frames), or three consecutive frames (preferably the previous, current, and next frames). Due to the well-known time domain alias cancellation (TDAC) associated with the MDCT, the three frame option achieves a perfect overlap of the input frames and is the most (possibly completely) accurate, at least for frames that contain time domain aliases. Obviously, the three frame inverse MDCT operates with a one frame lag. By allowing the use of an approximate time domain representation as input to the MDST transform, this delay can be avoided, and therefore the need to compensate for delays between different parts of the decoder system. In the two frame option, overlap / add enabling TDAC is performed on the first half of the frame, and aliases are only present in the second half. In the one frame option, there is no TDAC, so aliases occur throughout the entire frame. However, the MDST representation thus realised and used as a daytime signal with complex predictive coding can provide sufficient quality.
[0062] The decoding system shown in Fig. 3 can also operate in two non-predictive decoding modes. To decode a direct L / R coded stereo signal, the second switches 305, 310 are set in the lower position and the first switching assembly 302 is set in the pass-through mode. Thus, this signal is in L / R format upstream of the sum-and-difference stage 304, which converts it into M / S format, which is subjected to the inverse MDCT transformation and the final sum-and-difference operation. To decode a stereo signal provided in simultaneous M / S coding format, the first switching assembly 302 is set in the sum-and-difference mode, so that between the first switching assembly 302 and the sum-and-difference stage 304, the signal is in L / R format. The L / R format is more suitable than the M / S format from the point of view of TNS filtering. The processing downstream of the sum-and-difference stage 304 is the same as in the case of direct L / R decoding.
[0063] Figure 14 (14A to 14C) show three block diagrams of a decoder according to an embodiment of the present invention. Unlike the other block diagrams attached hereto, the connecting lines in Figure 14 show multi-channel signals. Specifically, the connecting lines are configured to transmit stereo signals having left / right, mid / side, downmix / residual, pseudo-left / pseudo-right channels, and other combinations.
[0064] Fig. 14A shows a decoder system for decoding a frequency domain representation of an input signal (for the purposes of this figure, shown as an MDCT representation). The decoder system is arranged to provide as its output a time domain representation of the stereo signal. This representation is generated on the basis of the input signal. To be able to decode an input signal coded by complex predictive stereo coding, the decoder system is provided with an upmix stage 1410. However, it is also possible to process input signals coded in other formats and possibly switching between several coding formats over time, for example an input signal in which a sequence of time frames coded by complex predictive coding is followed directly by a time portion coded by left / right coding. The functionality of the decoder system for processing different coding formats is realized by providing a connection line (pass-through) in parallel to said upmix stage 1410. A switch 1411 allows to select whether the output from the upmix stage 1410 (lower switch position in the figure) or the unprocessed signal obtained by the connection line (upper switch position in the figure) is to be fed to a decoder module arranged further downstream. In this embodiment, an inverse MDCT module 1412 is arranged downstream of the switch. The MDCT module 1412 converts the MDCT representation of the signal into a time domain representation. As an example, the signal fed to the upmix stage 1410 may be a stereo signal in downmix / residual format. The upmix stage 1410 is then applied to obtain the side signal and perform a sum / subtract operation to output left / right stereo signals (in the MDCT domain).
[0065] Figure 14B shows a decoder system similar to that shown in Figure 14A. The system is arranged to receive a bitstream as an input signal. The bitstream is first processed by a combined demultiplexer and inverse quantizer module 1420. This combined demultiplexer and inverse quantizer module 1420 provides an MDCT representation of the multi-channel stereo signal for further processing as a first output signal, as determined by the position of a switch 1422, which performs a similar function to the switch 1411 shown in Figure 14A. More precisely, the switch 1422 decides whether the first output from the demultiplexer and inverse quantizer is processed by the upmix stage 1421 and the inverse MDCT module 1423 (lower position) or by only the inverse MDCT module 1423 (upper position). The combined demultiplexer and inverse quantizer module 1420 also outputs control information. In this case, the control information relating to the stereo signal includes data indicating whether an up or down position of the switch 1422 is suitable for decoding the signal, or more abstractly, into which encoding format the stereo signal is to be decoded. The control information also includes parameters for adjusting the characteristics of the upmix stage, such as, for example, the value of the complex prediction coefficient α to be used in the complex predictive coding, as already explained.
[0066] FIG. 14C includes similar entities as shown in FIG. 14B, plus a first and a second frequency domain modification device 1431, 1435 located respectively upstream and downstream of the upmix stage 1433. For the purposes of this drawing, each frequency domain modification device is illustrated by a TNS filter. However, the term frequency domain modification device can also be understood to mean processes other than TNS filtering that can be applied before or after the upmix stage. Examples of frequency domain modifications include prediction, noise addition, bandwidth extension, non-linear processing. In some cases, psychoacoustic considerations and similar reasons, including the characteristics of the signal being processed and / or the settings of such frequency domain modification devices, make it advantageous to apply said frequency domain modification upstream of the upmix stage 1433 rather than downstream. In other cases, similar considerations make an upstream location of the frequency domain modification more preferable. Switches 1432, 1436 selectively activate the frequency domain modification devices 1431, 1435 in response to control information, such that the decoder system can select the desired configuration. As an example, Fig. 14C shows a configuration in which the stereo signal from the combined demultiplexer and inverse quantization module 1430 is first processed by the first frequency domain modification device 1431, then fed to the upmix stage 1433 and finally forwarded directly to the inverse MDCT module 1437, without passing through the second frequency domain adjustment device 1435. As explained in the Summary of the invention, this configuration is preferred over the option of performing TNS after the upmix in complex predictive coding.
[0067] II. Encoder system An encoder system according to the invention is described with reference to Fig. 5. Fig. 5 is a block diagram showing an encoder system for encoding a left / right (L / R) stereo signal as an output bitstream by complex predictive coding. The encoder system receives a time or frequency domain representation of the signal and supplies it to both a downmix stage and a prediction coefficient estimator. The real and imaginary parts of the prediction coefficients are supplied to the downmix stage to control the conversion of the left and right channels to the downmix and residual channels. The downmix and residual channels are then supplied to a final multiplexer MUX. If the signal was not supplied to the encoder as a frequency domain representation, it is converted to such a representation in the downmix stage or multiplexer.
[0068] One of the principles of predictive coding is to convert the left / right signals into a mid / side format, i.e.
number
number
[0069] Real part of prediction coefficient and α R and the imaginary part α I are quantized and / or coded jointly. However, preferably, the real and imaginary parts are quantized independently and uniformly, typically with a step size of 0.1 (a dimensionless number). According to the MPEG standard, the resolution of the frequency bands used for the complex prediction coefficients does not have to be the same as the resolution of the scale factor band (sfb, i.e. a set of MDCT lines using the same MDCT quantization step size and quantization range). In particular, the frequency band resolution of the prediction coefficients is psychoacoustically relevant, such as the Bark scale. Note that the resolution of the frequency bands changes when the transform length is changed.
[0070] As mentioned before, the encoder system according to the invention has the freedom to apply predictive stereo coding or not. The latter case implies a fallback to L / R or M / S coding. Such a decision can be made on a time frame or finer basis, or on a frequency band basis within a time frame. As mentioned above, the negative outcome of the decision is signalled to the decoding entity in various ways, for example by the value of a dedicated indicator bit in each frame, or by the presence or absence (or zero value) of the prediction coefficient values. Positive decisions are signalled as well. A particularly advantageous embodiment allows for a fallback without overhead, but makes use of a reserved fourth value of the 2-bit field ms_mask_present (MPEG-2 AAC, see ISO / IEC 131818-7 document), which is transmitted every time frame and is specified as follows: [Table 2] By redefining the value 11 to mean "complex predictive coding", the encoder can operate in all legacy modes, in particular in M / S and L / R coding modes, without compromising bitrate, and can advantageously receive a signal indicating complex predictive coding of a frame.
[0071] The actual decision may be based on a data rate vs. sound quality principle. As a measure of sound quality (as is often the case with available MDCT-based audio encoders), data obtained using a psychoacoustic model included in the encoder may be used. In particular, some encoder embodiments perform a rate-distortion optimized selection of the prediction coefficients. Thus, in such embodiments, the imaginary part, and possibly also the real part, of the prediction coefficients are set to zero if increasing the prediction gain does not save enough bits for the coding of the residual signal to justify the use of the bits required to code the prediction coefficients.
[0072] An embodiment of the encoder encodes TNS-related information into the bitstream. Such information includes the values of the TNS parameters used by the TNS (synthesis) filter at the decoder side. If both channels use the same set of TNS parameters, it is economical to include a signaling bit indicating that the parameters are the same, rather than transmitting the two sets of parameters separately. For example, information may also be included whether TNS should be applied before or after the upmix stage, based on a psychoacoustic evaluation of the two options.
[0073] As yet another optional feature, which is potentially beneficial from a complexity and bitrate point of view, the encoder is configured to use a separate limited bandwidth for the coding of the residual signal. The frequency bands above this limit are not transmitted to the decoder but are set to zero. In some cases, the energy content of the highest frequency bands is so small that they are already zero when quantized. Usual practice (see the max_sfb parameter in the MPEG standard) requires the use of the same bandwidth limit for both the downmix signal and the residual signal. Here, the inventors have empirically found that the residual signal has an energy content that is significantly more localized in low frequency bands than the downmix signal. Therefore, by imposing a dedicated bandwidth upper limit on the residual signal, a reduction in the bitrate is possible without a significant loss in sound quality. For example, this is achieved by transmitting two independent max_sfb parameters, one for the downmix signal and one for the residual signal.
[0074] It should be pointed out that although the issues of optimal determination of prediction coefficients, quantization and its coding, fallback to M / S or L / R mode, TNS filtering, upper bandwidth limit etc. have been described with reference to the decoder system shown in FIG. 5, the same are equally applicable to the embodiments disclosed in the embodiments described with reference to the subsequent figures.
[0075] Fig. 6 shows another encoder system according to the invention, arranged to perform complex predictive stereo coding. The system receives as input a time domain representation of a stereo signal, divided into successive, possibly overlapping, time frames, and including left and right channels. A sum-and-difference stage 601 converts this signal into a mid channel and a side channel. The mid channel is fed to both an MDCT module 602 and an MDST module 603, while the side channel is fed only to an MDCT module 604. A prediction coefficient drop 605 estimates the values of the complex prediction coefficients for each time frame, and possibly for individual frequency bands within a frame, as described above. The values of the coefficients α are fed as weights to weighted adders 606, 607, which construct a residual signal D as a linear combination of the MDCT and MDST representations of the mid signal and the MDCT representation of the side signal. The complex prediction coefficients are fed to the weighted adders 606, 607 preferably represented by the same quantization scheme used when it is encoded into the bitstream. This obviously provides a more faithful reconstruction since both the encoder and the decoder use the same values of the prediction coefficients. The residual signal, the mid signal (which is more appropriately called the downmix signal when it appears in combination with the residual signal), and the prediction coefficients are fed to a combined quantization and multiplexer stage 608, which encodes these signals and possibly further information into an output bitstream.
[0076] FIG. 7 shows a modification of the encoder shown in FIG. 6. As can be seen from the similar symbols in the figure, the encoder shown in FIG. 7 has a similar structure, but with the added functionality of operating in a direct L / R coding fallback mode. The encoder system is activated between the complex predictive coding mode and the fallback mode by a switch 710 located immediately upstream of the combined quantization and multiplexer stage 709. When the switch 710 is in the up position, the encoder operates in the fallback mode. The mid-side signal is fed to a sum-and-difference stage 705 from a point immediately downstream of the MDCT modules 702, 704. The sum-and-difference stage 705 converts the signal to a left / right signal and then sends it to the switch 710. The switch 710 connects the signal to the combined quantization and multiplexer stage 709.
[0077] Fig. 8 shows an encoder system according to the invention. In contrast to the encoder systems shown in Figs. 6 and 7, this embodiment obtains the MDST data required for complex predictive coding directly from the MDCT data, i.e. by a real-to-imaginary transformation in the frequency domain. The real-to-imaginary transformation can be any of the approaches described for the decoder systems of Figs. 2 and 4. It is important that the decoder calculation method matches the encoder calculation method to ensure faithful decoding. The same real-to-imaginary transformation method is used on the encoder and decoder sides. For the decoder embodiment, the part A with the real-to-imaginary transformation 804, enclosed by a dashed line, can be replaced by a similar variant or by using fewer input time frames. Similarly, the coding can be simplified by using any of the approximation approaches mentioned above.
[0078] At a high level, the encoder system of FIG. 8 has a different structure than would be obtained by simply replacing the MDST module of FIG. 7 by a (suitably connected) real-imaginary module. This architecture is clean and robust and computationally economical to implement the switch function between predictive coding and direct L / R coding. The input stereo signal is input to an MDCT transform module 801, which outputs a frequency domain representation of each channel. This is sent to both a final switch 808, which activates the encoder system between predictive and direct coding modes, and to a sum-and-difference stage 802. In direct L / R coding, or simultaneous M / S coding, which is performed in a time frame where the prediction coefficient α is set to zero, this embodiment only MDCT transforms, quantizes, and multiplexes the input signal. The latter two steps are performed by a combined quantization and multiplexer stage 807 located at the output of the system, which provides the bitstream. In predictive coding, each channel is further processed between the sum-and-difference stage 802 and the switch 808. A real-to-imaginary transform 804 derives MDST data from an MDCT representation of the mid signal and feeds it to both a prediction coefficient estimator 803 and a weighted summer 806. As in the encoder systems shown in Figures 6 and 7, a separate weighted summer 805 is used to combine the side signal with the weighted MDCT and MDST representation of the mid signal to form a residual channel signal, which is then encoded together with the mid (i.e. downmix) channel signal and the prediction coefficients by a combined quantization and multiplexing stage 807.
[0079] Now, with reference to Fig. 9, it will be explained that each embodiment of the encoder system can be combined with one or more TNS (analysis) filters. As mentioned before, it is often advantageous to apply TNS filtering to signals in downmix format. Thus, as shown in Fig. 9, adapting the encoder system of Fig. 7 to include TNS is done by adding a TNS filter 911 immediately upstream of the combined quantization and multiplexer stage 909.
[0080] Instead of the right / residual TNS filter 911b, two TNS filters (not shown) configured to process the right channel or the residual channel may be provided immediately upstream of the switch 910. In this way, each of the two TNS filters is always fed with a respective channel signal, allowing TNS filtering based on more time frames than just the current frame. As mentioned above, the TNS filter is an example of a frequency domain modification device, particularly one based on processing more frames than the current time frame, which benefits from such an arrangement as much as, or even more than, the TNS filter.
[0081] As another alternative to the embodiment shown in Fig. 9, the TNS filters for selective activation can be configured at one or more points for each channel. This is similar to the configuration of the decoder system shown in Fig. 4, where different sets of TNS filters can be connected by switches. This allows the most suitable stage for TNS filtering to be selected for each time frame. In particular, switching between different TNS locations is advantageous for switching between complex predictive stereo coding modes and other coding modes.
[0082] Figure 11 shows a variant on the encoder system of Figure 8, in which a second frequency domain representation of the downmix signal is obtained by a real-to-imaginary transform 1105. Like the decoder system shown in Figure 4, this encoder system also includes selectively activatable frequency domain modifier modules, one 1102 arranged upstream and one 1109 arranged downstream of the downmix stage. The frequency domain modules 1102, 1109, illustrated here by TNS filters, can be connected to the respective signal paths using four switches 1103a, 1103b, 1109a and 1109b.
[0083] III. Non-Device Embodiments Embodiments of the third and fourth aspects of the present invention are shown in figures 15 and 16. Figure 15 shows a method for decoding a bitstream into a stereo signal, comprising the steps of: 1. Input the bitstream. 2. Dequantizing the bitstream to obtain a first frequency domain representation of the downmix and residual channels of the stereo signal. 3. Compute a second frequency domain representation of the downmix channel. 4. Calculate the side channel signal based on the three frequency domain representations of the channel. 5. A stereo signal, preferably in left / right format, is calculated based on the side channels and the downmix channel. 6. The stereo signal thus obtained is output. Steps 3 to 5 may be considered as an upmixing process. Each of steps 1 to 6 is similar to the corresponding function of any of the decoder systems disclosed in the previous parts of this document, and implementation details can be found therein.
[0084] FIG. 16 illustrates a method for encoding a stereo signal into a bitstream signal, comprising the steps of: 1. Input a stereo signal. 2. Convert the stereo signal into a first frequency domain representation. 3. Determine the complex prediction coefficients. 4. Downmix the frequency domain representation. 5. Encode the downmix channel and the residual channel along with the complex prediction coefficients as a bitstream. 6. Output the bitstream. Each of steps 1 to 5 is similar to the corresponding function of any of the encoder systems disclosed in earlier parts of this document, and implementation details can be found therein.
[0085] Both methods can be expressed as computer readable instructions in the form of a software program and executed by a computer, the scope of protection of the invention extending to such software and to a computer program product for distributing such software.
[0086] IV. Experimental Evaluation The disclosed embodiments have been experimentally evaluated, and the most significant pieces of experimental material obtained in this process are summarized below.
[0087] The embodiment used in the experiments has the following features: (i) Each MDST spectrum (of a time frame) was calculated from the current, previous, and next MDCT spectra by two-dimensional finite impulse response filtering. (ii) The psychoacoustic model from the USAC stereo encoder was used. (iii) Instead of the PS parameters ICC, CLD and IPD, the real and imaginary parts of the complex prediction coefficient α were transmitted. The real and imaginary parts were processed separately, constrained to the range [-3.0, 3.0] and quantized with a step size of 0.1. They were time-differential coded and finally Huffman coded using the USAC scale factor codebook. The prediction coefficients were updated every other scale factor band, resulting in a frequency resolution similar to that of MPEG Surround (see, for example, ISO / IEC 23003-1). This quantization and coding scheme results in an average bit rate of about 2 kb / s for the stereo side information in a typical configuration with a target bit rate of 96 kb / s. (iv) The bitstream format was modified without breaking the current USAC bitstream since the 2-bit ms_mask_present bitstream element has only three possible values: a fourth value indicating complex prediction was used to allow a fallback mode for basic mid / side coding without wasting any bits (see the previous subsection of this disclosure for more information).
[0088] Listening tests were conducted using the MUSHRA method with eight test items played over headphones and a sampling rate of 48 kHz. Three, five or six subjects participated in each test.
[0089] The impact of different MDST approximations is evaluated to show the practical complexity vs. sound quality trade-off between these options. The results are shown in Fig. 12 and Fig. 13. The first shows the absolute scores obtained, while the latter shows the difference scores for 96s USAC cplf, i.e. for MDCT domain unified stereo coding with complex prediction using the current MDCT frame to compute an approximation of the MDST. It can be seen that the sound quality gain achieved by MDCT-based unified stereo coding increases when a computationally more complex approach is used to compute the MDST spectrum. Considering the average across all tests, the single frame-based system 96s USAC cplf significantly increases the coding efficiency with respect to conventional stereo coding. Similarly, even better results are obtained for 96s USAC cp3f, i.e. for MDCT domain unified stereo coding with complex prediction using the current, previous and next MDCT frames to compute the MDST.
[0090] V. Embodiments Furthermore, the present invention can be implemented as follows.
[0091] 1. A decoder system for decoding a bitstream signal into a stereo signal according to complex predictive stereo coding, comprising: an inverse quantization stage (202, 401) for providing a first frequency domain representation of the downmix signal (M) and the residual signal (D) based on said bitstream, each frequency domain representation having first spectral components representative of a spectral content of a corresponding signal expressed in a first subspace of a multidimensional space, said first spectral components being transform coefficients arranged in a time frame of transform coefficients, each block being generated by application of a transform to a time segment of the time domain signal; and a dequantization stage arranged downstream of the dequantization stage and configured to generate the stereo signal based on the downmix signal and the residual signal, said dequantization stage comprising: a module (206; 408) for calculating a second frequency domain representation of the downmix signal based on the first frequency domain representation of the downmix signal, the second frequency domain representation having second spectral components representative of a spectral content of the signal expressed in a second subspace of the multidimensional space including a portion of the multidimensional space not included in the first subspace, the module being configured to: determine a first intermediate component from the first spectral components; determine a second intermediate component by forming a combination of the first spectral components according to at least a portion of an impulse response; and determine the second spectral component from the second intermediate component; a weighted adder (210, 211; 406, 407) for calculating a side signal based on a first and a second frequency domain representation of the downmix signal encoded in the bitstream signal, a first frequency domain representation of the residual signal, and complex prediction coefficients (α); and The downmix signal is coupled to a first frequency domain representation of the side signal, the downmix signal being coupled to a first frequency domain representation of the side signal, and the stereo signal is coupled to a first frequency domain representation of the side signal.
[0092] Furthermore, the present invention can be implemented as follows: A decoder system for decoding a bitstream signal into a stereo signal by complex predictive stereo coding, comprising: an inverse quantization stage (301) for providing first frequency domain representations of a downmix signal (M) and a residual signal (D) based on said bitstream signal, each of said first frequency domain representations having first spectral components representative of the spectral content of a corresponding signal expressed in a first subspace of a multidimensional space; and a dequantization stage arranged downstream of the dequantization stage and configured to generate the stereo signal based on the downmix signal and the residual signal, said dequantization stage comprising: a module (306, 307) for calculating a second frequency domain representation of the downmix signal based on a first frequency domain representation of the downmix signal, the second frequency domain representation having a spectral content of the signal expressed in a second subspace of the multidimensional space that includes a part of the multidimensional space that is not included in the first subspace, comprising: an inverse transform step (306) for calculating a time domain representation of the downmix signal based on the first frequency domain representation of the downmix signal in the first subspace of the multidimensional space; and a transform step (307) for calculating a second frequency domain representation of the downmix signal based on the time domain representation of the signal; a weighted adder (308, 309) for calculating a side signal based on a first and a second frequency domain representation of the downmix signal encoded in the bitstream signal, a first frequency domain representation of the residual signal, and complex prediction coefficients (α); and The downmix signal is coupled to a first frequency domain representation of the side signal, the downmix signal being coupled to a first frequency domain representation of the side signal, and the upmix signal is coupled to a first frequency domain representation of the side signal, the upmix signal being coupled to a first frequency domain representation of the side signal, and
[0093] The invention can also be implemented as follows: A decoder system having the features of the independent decoder system claim, the module for calculating the second frequency domain representation of the downmix signal comprising: an inverse transform step (306) of calculating a time domain representation of the downmix signal and / or the chroma signal based on a first frequency domain representation of each signal in a first subspace of the multidimensional space; and a transform step (307) for calculating a second frequency domain representation of each signal based on the time domain representation of the signal; Preferably, the inverse transform step (306) performs an inverse modified discrete cosine transform and the transform step performs a modified discrete cosine transform.
[0094] In the above decoder system, the stereo signal may be represented in the time domain and the decoder system may further comprise: a switching assembly (302) disposed between the inverse quantization stage and the upmix stage, the switching assembly being capable of functioning either as (a) a pass-through stage for use in simultaneous stereo encoding; or (b) a sum-and-difference stage for use in direct stereo encoding; a further inverse transform stage (311) disposed in the upmix stage, which calculates a time domain representation of the side signal; (a) a further sum / difference stage (304) connected to a point downstream of the switching assembly (302) and upstream of the upmix stage; or (b) a selector device (305, 310) arranged upstream of the inverse transformation stage (306, 301) configured to be selectively connected to either the downmix signal resulting from the switching assembly (302) or the side signal resulting from the weighted adder (308, 309).
[0095] VI. Conclusion Further embodiments of the present invention will be apparent to those skilled in the art upon reading the above description. Although the present specification and drawings disclose embodiments and examples, the present invention is not limited to these specific examples. Numerous modifications and variations can be made without departing from the scope of the present invention, which is defined in the appended claims.
[0096] It is to be noted that the method and apparatus disclosed in this application can be applied to the coding of signals having more than two channels with appropriate modifications within the ability of a person skilled in the art, including routine experimentation. It is emphasized that the signals, parameters and matrices mentioned in connection with the described embodiments may be frequency-varying or frequency-invariant and / or time-varying or time-invariant. The described computational steps may be performed frequency-by-frequency or for all frequencies at once, and all entities may be implemented with frequency-selective operation. For the purposes of the application, any quantization scheme may be adapted according to a psychoacoustic model. It is further noted that the various sum-difference transforms, i.e., the transform from downmix / residual format to pseudo-L / R format and the L / R-to-M / S transform and the M / S-to-L / R transform, all have the following forms:
number
[0097] The systems and methods disclosed herein may be implemented as software, firmware, hardware, or a combination thereof. Some or all of the components may be implemented as software executed by a digital signal processor or microprocessor, or may be implemented as hardware or special purpose integrated circuits. Such software may be distributed on a computer readable medium. Computer readable media includes computer storage media and communication media. As known to those skilled in the art, computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory and other memory technologies, CD-ROM, digital versatile disks (DVDs) and other optical disk storage media, magnetic cassettes, magnetic tapes, magnetic disk storage and other magnetic storage devices, or any other medium capable of storing the desired information. Additionally, as known to those skilled in the art, communication media generally embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and includes any information delivery media. In addition, the following note is added: (Supplementary Note 1) A decoder system for providing a stereo signal by complex predictive stereo coding, comprising: an upmix stage configured to generate the stereo signal based on first frequency domain representations of a downmix signal and a residual signal, each first frequency domain representation having first spectral components representative of a spectral content of a corresponding signal expressed in a first subspace of a multidimensional space, the upmix stage comprising: a module for calculating a second frequency domain representation of the downmix signal based on a first frequency domain representation of the downmix signal, the second frequency domain representation having second spectral components representative of a spectral content of the signal expressed in a second subspace of the multidimensional space, the second subspace including a part of the multidimensional space that is not included in the first subspace; a weighted adder for calculating a side signal based on first and second frequency domain representations of the downmix signal encoded in the bitstream signal, a first frequency domain representation of the residual signal, and complex prediction coefficients; an upmix stage comprising a sum / difference stage for calculating the stereo signal based on the downmix signal and a first frequency domain representation of the side signal, The upmix stage is further operable in a pass-through mode, in which the downmix signal and the residual signal are fed directly to the sum-and-difference stage. (Supplementary Note 2) The downmix signal and the residual signal are segmented into time frames, and the upmix stage is configured to receive, for each time frame, a 2-bit data field associated with that frame and to operate in an active mode or a pass-through mode depending on the value of the data field. 2. A decoder system as claimed in claim 1. (Supplementary Note 3) The downmix signal and the residual signal are segmented into time frames; the upmix stage is further adapted to receive, for each time frame in the MPEG bitstream, a ms_mask_present field associated with that frame, and to operate in an active mode or a pass-through mode depending on the value of the ms_mask_present field. 2. A decoder system as claimed in claim 1. (Supplementary Note 4) The method further comprises an inverse quantization stage arranged upstream of the upmix stage, the inverse quantization stage providing the first frequency domain representation of the downmix signal and the residual signal based on a bitstream signal. 4. A decoding system according to any one of claims 1 to 3. (Supplementary Note 5) The first spectral component has real-valued values represented in the first subspace; the second spectral component has an imaginary value represented in the second subspace; Optionally, the first spectral component is determined by one of a discrete cosine transform (DCT) or a modified discrete cosine transform (MDCT); Optionally, the second spectral component is determined by one of a discrete sine transform (DST) or a modified discrete sine transform (MDST). 5. A decoder system according to any one of claims 1 to 4. (Supplementary Note 6) At least one temporal noise shaping TNS module arranged upstream of the upmix stage; at least one further TNS module arranged downstream of said upmix stage; a selector device for selectively activating either (a) the TNS module upstream of the upmix stage, or (b) the further TNS module downstream of the upmix stage. 6. A decoder system according to any one of claims 1 to 5. (Supplementary Note 7) The downmix signal is partitioned into consecutive time frames, each time frame being associated with a value of a complex prediction coefficient; the module for computing a second frequency domain representation of the downmix signal is configured to deactivate itself in response to an absolute value of the imaginary part of the complex prediction coefficients being smaller than a predetermined tolerance value for a time frame so as to not generate an output for that time frame. 6. A decoder system as claimed in claim 5. (Supplementary Note 8) The downmix signal time frame is further partitioned into frequency bands, each frequency band being associated with a value of the complex prediction coefficient; the module for computing a second frequency domain representation of the downmix signal is configured to deactivate itself in response to an absolute value of the imaginary part of the complex prediction coefficients being smaller than a predetermined tolerance value for a frequency band of a time frame so as to not generate an output for that frequency band. 8. A decoder system as claimed in claim 7. (Supplementary Note 9) The first spectral components are transform coefficients arranged in a time frame of transform coefficients, each block being generated by application of a transform to a time segment of a time domain signal; a module for computing a second frequency domain representation of the downmix signal comprising: determining a first intermediate component from the first spectral components; forming a combination of the first spectral components according to at least a portion of an impulse response to determine a second intermediate component; and determining a second spectral component from the second intermediate component. 9. A decoder system according to any one of claims 1 to 8. (Supplementary Note 10) A part of the impulse response is based on the frequency response characteristic of the transformation, Optionally, a frequency response characteristic of the transformation depends on a characteristic of an analysis window function applied to the transformation of the time segments of the signal. 10. A decoder system as claimed in claim 9. (Supplementary Note 11) The module for calculating a second frequency domain representation of the downmix signal comprises: (a) Simultaneous time frames of a first spectral component; (b) the contemporaneous and previous time frames of the first spectral component; and (c) determining each time frame of the second spectral component based on one of a concurrent, a previous, and a later time frame of the first spectral component; 11. A decoder system according to claim 9 or 10. (Supplementary Note 12) The module for calculating the second frequency domain representation of the downmix signal is configured to calculate an approximate second spectral representation having an approximate second spectral component determined by a combination of at least two time-adjacent and / or frequency-adjacent first spectral components. 12. A decoder system according to any one of claims 1 to 11. (Supplementary Note 13) The stereo signal is represented in the time domain, and the decoder system further comprises: a switching assembly disposed between said inverse quantization stage and said upmix stage, said switching assembly being capable of functioning as either (a) a pass-through stage, or (b) a sum-and-difference stage, thereby being capable of switching between direct and simultaneously encoded stereo input signals; an inverse transform stage configured to calculate a time domain representation of said stereo signal; a selector device arranged upstream of the inverse transform stage and configured to selectively connect it to either (a) a point downstream of the upmix stage at which a stereo signal obtained by complex prediction is fed to the inverse transform stage, or (b) a point downstream of the switching assembly and upstream of the upmix stage at which a stereo signal obtained by direct stereo coding is fed to the inverse transform stage, 13. A decoder system according to any one of claims 1 to 12. (Supplementary Note 14) An encoder system for encoding a stereo signal using complex prediction as a signal having a downmix channel, a residual channel, and complex prediction coefficients, comprising: an estimator for estimating complex prediction coefficients; (a) converting the stereo signal into a frequency domain representation of a downmix signal and a residual signal having a relationship determined by values of the complex prediction coefficients; and (b) an encoding stage operable to act as a pass-through stage and to feed the stereo signal to be encoded directly to a multiplexer. (Supplementary Note 15) configured to encode a stereo signal by a bitstream signal by complex predictive stereo coding, a multiplexer for receiving outputs from the encoding stage and the estimator and encoding the outputs according to the bitstream signal; 15. The encoder system of claim 14. (Supplementary Note 16) The estimator determines the complex prediction coefficients by minimizing the power of the residual signal over time or the average power of the residual signal. 16. The encoder system of claim 14 or 15. (Supplementary Note 17) The stereo signal has a downmix channel and a side channel; the encoding stage is configured to receive a first frequency domain representation of the stereo signal, the first frequency domain representation having first spectral components representing a spectral content of the corresponding signal expressed in a first subspace of a multidimensional space; The encoding step further comprises: a module for calculating a second frequency domain representation of the downmix channel based on a first frequency domain representation of the downmix signal, the second frequency domain representation having second spectral components representative of a spectral content of the signal expressed in a second subspace of the multidimensional space, the second subspace including a part of the multidimensional space that is not included in the first subspace; a weighted adder for calculating a residual signal based on the first and second frequency domain representations of the downmix channel, the first frequency domain representation of the side channel and the complex prediction coefficients, the estimator receives the downmix channel and a side channel and determines the complex prediction coefficients in order to minimize the power of the residual signal over time or to minimize the average power of the residual signal. 17. An encoder system according to any one of claims 14 to 16. (Supplementary Note 18) The encoding step comprises: a sum / subtract stage for converting the stereo signal into a jointly encoded stereo signal having a downmix channel and a side channel; a transforming step for providing an oversampled frequency domain representation of the downmix channel and a critically sampled frequency domain representation of the side channel, the oversampled frequency domain representation preferably having complex spectral components; a weighted adder for calculating a residual signal based on the oversampled frequency domain representation of the downmix channel, the critically sampled frequency domain representation of the side channel and the complex prediction coefficients, the estimator receives the residual signal and determines the complex prediction coefficients to minimize the power of the residual signal or to minimize the average power of the residual signal; Preferably, the transform stage comprises a Modified Discrete Cosine Transform stage arranged in parallel with a Modified Discrete Sine Transform stage MDST, which together provide the oversampled frequency domain representation of the downmix channel. 17. An encoder system according to any one of claims 14 to 16. (Supplementary Note 19) A decoding method for providing a stereo signal by complex predictive stereo coding, comprising the steps of: receiving first frequency domain representations of the downmix signal and the residual signal, each of the first frequency domain representations having first spectral components representative of a spectral content of the corresponding signal expressed in a first subspace of a multidimensional space; receiving a control signal; Depending on the value of the control signal, (a) upmixing the downmix signal and a residual signal using an upmix stage to obtain the stereo signal, - calculating a second frequency domain representation of the downmix signal based on the first frequency domain representation of the downmix signal, the second frequency domain representation having second spectral components representative of a spectral content of the signal expressed in a second subspace of the multidimensional space, the second subspace comprising a part of the multidimensional space that is not included in the first subspace; calculating a side signal based on first and second frequency domain representations of the downmix signal encoded in the bitstream signal, a first frequency domain representation of the residual signal, and complex prediction coefficients; and calculating the stereo signal by applying a first frequency domain representation of the downmix signal and the side signal to a sum-difference transform. (b) interrupting the upmixing step. (Supplementary Note 20) The first spectral component has real-valued values represented in the first subspace; the second spectral component has an imaginary value represented in the second subspace; Optionally, the first spectral component is determined by one of a discrete cosine transform (DCT) or a modified discrete cosine transform (MDCT); Optionally, the second spectral component is determined by one of a discrete sine transform (DST) or a modified discrete sine transform (MDST). 20. A decryption method as described in claim 19. (Supplementary Note 21) The downmix signal is partitioned into consecutive time frames, each time frame being associated with a value of a complex prediction coefficient; the step of calculating a second frequency domain representation of the downmix signal is interrupted in response to the absolute value of the imaginary part of the complex prediction coefficients being smaller than a predetermined tolerance for a time frame, such that no output is generated for that time frame. 21. A decryption method as described in claim 20. (Supplementary Note 22) The downmix signal time frame is further partitioned into frequency bands, each frequency band being associated with a value of the complex prediction coefficient; the step of calculating a second frequency domain representation of the downmix signal is interrupted in response to an absolute value of the imaginary part of the complex prediction coefficients being smaller than a predetermined tolerance for a frequency band of a time frame, such that no output is generated for that frequency band. 22. A decryption method according to claim 21. (Supplementary Note 23) The first spectral components are transform coefficients arranged in a time frame of transform coefficients, each block being generated by application of a transform to a time segment of a time domain signal; The step of calculating a second frequency domain representation of the downmix signal comprises: determining a first intermediate component from the first spectral component; forming a combination of the first spectral components according to at least a portion of an impulse response to obtain a second intermediate component; and determining the second spectral component from the second intermediate component. 21. A decryption method as described in claim 20. (Appendix 24) A part of the impulse response is based on the frequency response characteristic of the transformation, Optionally, a frequency response characteristic of the transformation depends on a characteristic of an analysis window function applied to the transformation of the time segments of the signal. 24. A decryption method as described in claim 23. (Supplementary Note 25) The step of calculating the second frequency domain representation includes obtaining each time frame of the second spectral component using as input one of (a) the contemporaneous time frame of the first spectral component, (b) the contemporaneous and previous time frames of the first spectral component, and (c) the contemporaneous, previous, and later time frames of the first spectral component. 25. A decryption method as described in claim 24. (Supplementary Note 26) The step of calculating the second frequency domain representation of the downmix signal includes a step of calculating an approximate second spectral representation having an approximate second spectral component determined by a combination of at least two time-adjacent and / or frequency-adjacent first spectral components. 26. A decoding method according to any one of claims 19 to 25. (Supplementary Note 27) The stereo signal is represented in the time domain, and the method further comprises: - omitting said upmixing step depending on whether said bitstream signal is encoded by direct stereo encoding or by simultaneous stereo encoding; inverse transforming the bitstream signal to obtain the stereo signal. 27. A decoding method according to any one of claims 19 to 26. (Supplementary Note 28) The method according to the present invention further comprises the steps of: omitting the steps of transmitting the time-domain representation of the downmix signal and calculating a side signal depending on whether the bitstream is encoded by direct stereo encoding or simultaneous stereo encoding; and inverse transforming a frequency domain representation of each channel encoded by the bitstream signal to obtain the stereo signal. 28. A decryption method as described in claim 27. (Supplementary Note 29) A method for encoding a stereo signal using a bitstream by complex predictive stereo coding, comprising the steps of: determining complex prediction coefficients; transforming the stereo signal into a first frequency domain representation of a downmix signal and a residual signal having a relationship determined by the complex prediction coefficients, the first frequency domain representation having first spectral components representing the spectral content of corresponding signals expressed in a first subspace of a multi-dimensional space; encoding the downmix channel, the residual channel and the complex prediction coefficients as the bitstream. (Supplementary Note 30) The step of determining the complex prediction coefficients is performed to minimize the power of the residual signal or the average power of the residual signal over time. 29. The encoding method of claim 29. (Supplementary Note 31) A step of defining or recognizing partitions of the stereo signal into time frames; and for each time segment, encoding or selecting the stereo signal in this time segment by at least one of the following options: direct stereo encoding, joint stereo encoding, and complex predictive stereo encoding; if direct stereo encoding is selected, the stereo signal is converted to a frequency domain representation of a left channel and a right channel and encoded into the bitstream; If simultaneous stereo encoding is selected, the stereo signal is converted to a frequency domain representation of a downmix channel and a side channel and encoded into the bitstream. 31. The encoding method of claim 29 or 30. (Appendix 32) The option that provides the highest sound quality according to a given psychoacoustic model is selected; 32. The encoding method of claim 31. (Supplementary Note 33) The method further comprises the step of defining or recognizing partitions of the stereo signal into time frames; the stereo signal has a downmix channel and a side channel; Transforming the stereo signal into a first frequency domain representation of a downmix channel and a residual channel comprises: - calculating a second frequency domain representation of the downmix signal based on the first frequency domain representation of the downmix channel, the second frequency domain representation having second spectral components representative of a spectral content of the signal expressed in a second subspace of the multidimensional space, the second subspace including a part of the multidimensional space that is not included in the first subspace; constructing a residual signal based on a first and a second frequency domain representation of the downmix channel, a first frequency domain representation of the side channel and the complex prediction coefficients; said step of determining complex prediction coefficients is performed one time frame at a time by minimizing an average power of a residual signal in each time frame; 31. The encoding method of claim 29 or 30. (Supplementary Note 34) A step of converting the stereo signal into a jointly encoded stereo signal having a downmix channel and a side channel; - converting said downmix channels into an oversampled frequency domain representation, preferably having complex spectral content; converting the side channel to a critically sampled, preferably real-valued, frequency domain representation; and calculating a residual signal based on the oversampled frequency domain representation of the downmix channel, the critically sampled frequency domain representation of the side channel and the complex prediction coefficients, the determination of the complex prediction coefficients is performed by feedback control on the residual signal thus calculated in order to minimize its power or its average power, 34. The encoding method according to any one of claims 29 to 33. (Supplementary Note 35) The conversion of the downmix channels to an oversampled frequency domain representation is performed by applying MDCT and MDST and concatenating the outputs. 35. The encoding method of claim 34. (Supplementary Note 36) A computer program product having a computer-readable medium storing instructions which, when executed by a general purpose computer, perform the method of any one of Supplements 19 to 35.< / nd>
Claims
1. 1. An apparatus for outputting a stereo audio signal having a left channel and a right channel, the apparatus comprising: a demultiplexer for receiving an audio bitstream and decoding at least one prediction coefficient from the audio bitstream, the audio bitstream being segmented into frames, and a value of the at least one prediction coefficient may change for each frame; a decoder configured to generate a downmix signal and a residual signal from the audio bitstream; an upmixer configured to operate in a predictive mode or a non-predictive mode and to output the left channel and the right channel as the stereo audio signal; when the upmixer operates in the predictive mode, the residual signal represents a difference between a side signal and a prediction of the side signal, and the upmixer generates the left channel and the right channel from a combination of the downmix signal, the residual signal and the at least one prediction coefficient; when the upmixer operates in the non-predictive mode, the residual signal represents the side signal, and the upmixer generates the left channel based on a sum of the downmix signal and the residual signal and generates the right channel based on a difference between the downmix signal and the residual signal. Device.
2. the at least one prediction coefficient reduces or minimizes the energy of the residual signal.
2. The apparatus of claim 1.
3. a noise shaper configured to shape noise associated with the downmix signal, the noise shaper being arranged upstream of the upmixer.
2. The apparatus of claim 1.
4. the noise shaper is a temporal noise shaper configured to shape the noise over time; 4. The apparatus of claim 3.
5. and when the upmixer operates in the predictive mode, the upmixer generates the left and right channels using a filter having three taps.
2. The apparatus of claim 1.
6. the downmix signal includes a mid signal formed by a linear combination of an original left channel and an original right channel; 2. The apparatus of claim 1.
7. the at least one prediction coefficient is a real-valued coefficient.
2. The apparatus of claim 1.
8. the at least one prediction coefficient is a complex-valued coefficient.
2. The apparatus of claim 1.
9. the upmixer combines the side signal with the downmix signal by adding a version of the downmix signal to a version of the side signal to generate the left channel and subtracting the version of the side signal from the version of the downmix signal to generate the right channel.
2. The apparatus of claim 1.
10. The apparatus of claim 1 , wherein the upmixer is configured, when operating in the predictive mode, to add the residual signal to the side signal.
11. the downmix signal is divided into frequency bands, the at least one prediction coefficient comprises a prediction coefficient for each frequency band, and the demultiplexer is configured to receive the audio bitstream and to decode each of the prediction coefficients from the audio bitstream; and when the upmixer operates in the predictive mode, the upmixer generates the left channel and the right channel from a combination of the downmix signal, the residual signal, and each of the prediction coefficients.
2. The apparatus of claim 1.
12. 1. A method for outputting a stereo audio signal having a left channel and a right channel, the method comprising: receiving an audio bitstream and decoding at least one prediction coefficient from the audio bitstream, the audio bitstream being segmented into frames, and a value of the at least one prediction coefficient may change for each frame; generating, in a decoder, a downmix signal and a residual signal from the audio bitstream; upmixing in a predictive or non-predictive mode; outputting the left channel and the right channel as the stereo audio signal; when the upmixing operates in the predictive mode, the residual signal represents a difference between a side signal and a prediction of the side signal, and the upmixing generates the left channel and the right channel from a combination of the downmix signal, the residual signal and the at least one prediction coefficient; When the upmixing operates in the non-predictive mode, the residual signal represents the side signal, and the upmixing generates the left channel based on a sum of the residual signal and the downmix signal passed through the decoder, and generates the right channel based on a difference between the residual signal and the downmix signal. method.
13. A non-transitory computer readable medium comprising instructions that, when executed by a processor, perform the method of claim 12.
Citation Information
Patent Citations
signal processing
JP2005521921A
Methods for Improving Performance of Prediction-Based Multi-Channel Reconstruction
JP2008517337A
Enhanced coding and parameterization in multi-channel downmixed object coding
JP2010507115A
A parametric stereo upmix apparatus, a parametric stereo decoder, a parametric stereo downmix apparatus, a parametric stereo encoder
WO2009141775A1