Decoder system, decoding method, and computer program

The complex prediction stereo encoding method addresses inefficiencies in high-bitrate stereo audio coding by using frequency domain representations and adaptive decoding, achieving efficient and flexible encoding with reduced complexity and improved audio quality.

JP7703123B2Active Publication Date: 2025-07-04DOLBY INTERNATIONAL AB
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025039836
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2010-04-09
Filing Date
2025-03-13
Publication Date
2025-07-04
Estimated Expiration
2031-04-06

AI Technical Summary

Technical Problem

Existing stereo audio coding methods, particularly at high bitrates, face inefficiencies in computational complexity and delay due to the use of complex-valued upmix matrices and QMF filter banks, which are not optimal for high bitrates.

Method used

A method and apparatus for high-efficiency stereo encoding using complex prediction stereo encoding, involving an upmix stage that generates stereo signals based on downmix and residual signals in the frequency domain, with a module calculating a second frequency domain representation and a weighted adder using complex prediction coefficients, allowing for adaptive decoding between direct and complex prediction coding.

Benefits of technology

Preserves the advantages of unified stereo coding at high bitrates with reduced computational complexity, enabling faithful signal reconstruction and flexibility in encoding modes, maintaining or improving audio quality without significant increases in complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007703123000014
    Figure 0007703123000014
  • Figure 0007703123000015
    Figure 0007703123000015
  • Figure 0007703123000016
    Figure 0007703123000016
Patent Text Reader

Abstract

To provide methods and apparatus for stereo coding that are computationally efficient also in high bit-rate range.SOLUTION: The invention provides methods and devices for stereo encoding and decoding using complex prediction in a frequency domain. In one embodiment, a decoding method, for obtaining an output stereo signal from an input stereo signal encoded by complex prediction coding and having first frequency-domain representations representing two input channels, includes the upmixing steps of: (i) computing a second frequency-domain representation of a first input channel; and (ii) computing an output channel on the basis of the first and second frequency-domain representations of the first input channel, the first frequency-domain representation of the second input channel, and a complex prediction coefficient. The upmixing can be suspended responsive to control data.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention disclosed herein generally relates to stereo audio coding, and more particularly to a stereo coding technique using complex prediction in the frequency domain.

Background Art

[0002] By jointly coding the left (L) and right (R) channels of a stereo signal, coding becomes more efficient compared to coding the L and R independently. A common approach to joint stereo coding is mid / side (M / S) coding. Here, the mid (M) signal is formed by adding the L and R signals. For example, the M signal is

Number

Number

[0003] In the AAC (Advanced Audio Coding) standard of MPEG (Moving Picture Experts Group) (refer to the standard document ISO / IEC 13818-7), it is time- and frequency-variable, and L / R stereo coding and M / S stereo coding can be selected. In this way, the stereo encoder can apply L / R coding to a certain frequency band of the stereo signal, and M / S coding is used for encoding other frequency bands of the stereo signal (frequency-variable). Furthermore, the encoder can switch between L / R coding and M / S coding temporally (time-variable). In MPEG AAC, stereo encoding is performed in the frequency domain, more specifically, in the MDCT (Modified Discrete Cosine Transform) domain. This enables the adaptive selection of either L / R coding or M / S coding in a frequency- and time-variable manner.

[0004] Parametric stereo coding is a method for efficiently encoding a stereo audio signal as a monoral signal and a small amount of side information serving as stereo parameters. This is part of the MPEG-4 audio standard (refer to the standard document ISO / IEC14496-3). The monoral signal can be encoded using any audio encoder. Since the stereo parameters are incorporated into an appendix of the mono bit stream, it is completely forward and backward compatible. At the decoder, the monoral signal is decoded first, and then the stereo signal is reconstructed using the stereo parameters. The decorrelated version of the decoded mono signal has a zero cross-correlation with the mono signal. This decorrelated signal is generated by a decorrelator, for example, by an appropriate all-pass filter including a delay line. Basically, the decorrelated signal has the same spectral and temporal energy distribution as the mono signal. The monoral signal is input into the upmix process together with the decorrelated signal. This process is controlled by the stereo parameters to reconstruct the stereo signal. For more detailed information, refer to Non-Patent Document 1.

[0005] MPEG Surround (MPS; see ISO / IEC 23003-1 and Non-Patent Document 2) combines the principle of parametric stereo coding with the principle of residual coding, replacing the uncorrelated signal with the transmitted residual to improve the perceived audio quality. Residual coding is performed by downmixing the multi-channel signal and optionally extracting a spatial cue. In the downmixing process, a residual signal representing the error signal is calculated, encoded, and transmitted. The residual signal replaces the uncorrelated signal at the decoder. In the hybrid approach, the residual signal replaces the uncorrelated signal in a certain frequency band, preferably in a relatively low band.

[0006] The current MPEG Unified Speech and Audio Coding (USAC) system, although two examples are shown in FIG. 1, has a complex-valued quadrature mirror filter (QMF) bank located downstream of the core decoder. The QMF representation obtained as the output of this filter bank is complex-valued and is therefore oversampled by a factor of two and can be configured as a downmix signal (i.e., a mid signal) M and a residual signal D. An upmix matrix with complex-valued components can be used for this. The L and R signals (in the QMF region) are

Number

[0007] The above encoding configuration is generally well-suited for low bitrates, generally less than 80 kb / s, but is not optimal in terms of computational complexity for high bitrates. More specifically, at high bitrates, generally the SBR tool is not used (since it does not improve the encoding efficiency). Next, in a decoder without an SBR stage, a complex-valued upmix matrix results in the use of a QMF filter bank, which is computationally intensive and causes delay (for a frame length of 1024 samples, a delay of 961 samples is caused by the QMF analysis / synthesis filter bank). This clearly indicates the need for a more efficient encoding configuration.

Prior Art Documents

Non-Patent Documents

[0008]

Non-Patent Document 1

Summary of the Invention

Means for Solving the Problems

[0009] One object of the present invention is to provide a method and apparatus for high-efficiency stereo encoding even in a high bit rate range.

[0010] The present invention achieves this object by providing an encoder and a decoder, a coding and decoding method, and a computer program product for encoding and decoding, respectively, as defined in the independent claims. The dependent claims define embodiments of the present invention.

[0011] In a first aspect, the present invention provides the following system. That is, a decoder system that provides a stereo signal by complex prediction stereo encoding: An upmix stage configured to generate the stereo signal based on a first frequency domain representation of a downmix signal (M) and a residual signal (D), each first frequency domain representation having a first spectral component representing the spectral content of the corresponding signal represented in a first subspace of a multi-dimensional space, the upmix stage comprising: A module that calculates a second frequency domain representation of the downmix signal based on the first frequency domain representation of the downmix signal, the second frequency domain representation having a second spectral component representing the spectral content of a signal represented in a second subspace of the multi-dimensional space, the second subspace including a part of the multi-dimensional space not included in the first subspace, the module; An upmix stage having a weighted adder that calculates a side signal (S) based on the first and second frequency domain representations of the downmix signal encoded in the bitstream signal, the first frequency domain representation of the residual signal, and a complex prediction coefficient (α); having a sum / difference stage that calculates the stereo signal based on a first frequency domain representation of the downmix signal and the side signal, The upmix stage is further operable in a pass-through mode in which the downmix signal and the residual signal are directly supplied to the sum / difference stage.

[0012] In a second aspect, the present invention provides the following system. That is, An encoder system that encodes a stereo signal into a bitstream signal by complex prediction stereo coding, an estimator that estimates complex prediction coefficients, (a) a coding stage operable to convert the stereo signal into a frequency domain representation of a downmix signal and a residual signal having a relationship determined by the value of the complex prediction coefficients, and a multiplexer that receives outputs from the coding stage and the estimator and encodes this into the bitstream signal.

[0013] In the third and fourth aspects of the present invention, a method for encoding a stereo signal into a bitstream and a method for decoding a bitstream into at least one stereo signal are provided. The technical features of each method are the same as the technical features of the encoder system and the decoder system, respectively. In the fifth and sixth aspects, the present invention provides a computer program product including instructions for executing each method on a computer.

[0014] The present invention benefits from the advantages of unified stereo coding in the MPEG USAC system. These advantages are preserved even at high bitrates without significantly increasing the computational complexity associated with the QMF-based approach, where SBR is not generally utilized. This is possible because the critically sampled MDCT transform, which is fundamental to the MPEG USAC transform, has the same downmix and residual channel coded audio bands, and at least when the upmix process does not involve decorrelation, according to the present invention, it can also be used in complex prediction stereo coding. This means that an additional QMF transform is no longer necessary. Representative embodiments of complex prediction stereo coding in the QMF domain significantly increase the number of arithmetic operations per unit time compared to conventional L / R or M / S stereo. Therefore, the encoding device according to the present invention seems to be competitive at such bitrates due to its modest computational load and the provision of high audio quality.

[0015] As those skilled in the art will notice, due to the fact that the upmix stage can also operate in a pass-through mode, the decoder can be adaptively decoded by conventional direct coding or simultaneous coding, and complex prediction coding, depending on the decision on the encoder side. Thus, it can be guaranteed that at least the same level is maintained when the decoder cannot raise the audio quality level more aggressively than conventional direct L / R stereo coding or simultaneous M / S stereo coding. Therefore, the decoder according to this aspect of the present invention can be regarded as a superset with respect to the background art from a functional perspective.

[0016] As an advantage over QMF-based predictive coding stereo, (except for quantization errors that can be arbitrarily small), complete reconstruction of the signal is possible.

[0017] Thus, the present invention provides an encoding apparatus for performing conversion-based stereo encoding by complex prediction. Preferably, the apparatus according to the present invention is not limited to complex prediction stereo encoding, and can also operate with direct L / R stereo encoding or simultaneous M / S encoding according to the background art, and can select the most suitable encoding method for a specific application or a specific time.

[0018] The oversampled representation of the signal (e.g., complex representation) includes both the first and second spectral components and is used as the basis for complex prediction according to the present invention. Therefore, a module for calculating such an oversampled representation is configured in the encoder system and decoder system according to the present invention. The spectral components refer to the first and second subspaces of the multi-dimensional space. This is a set of time-dependent functions of a given time length (e.g., the length of a predetermined time frame) sampled at a finite sampling frequency. It is well known that functions in this multi-dimensional space can be approximated by a finite weighted sum of basis functions.

[0019] As will be apparent to those skilled in the art, an encoder configured to cooperate with a decoder is provided with an equivalent module that provides an oversampled representation that is the basis of predictive encoding so as to enable faithful reproduction of the encoded signal. Such an equivalent module is the same or a similar module, or a module having the same or similar transfer characteristics. In particular, the encoder and decoder modules may be similar or dissimilar units that execute computer programs that perform equivalent mathematical operations, respectively.

[0020] In certain embodiments of the decoder system and encoder system, the first spectral component has a real value represented in the first subspace, and the second spectral component has an imaginary value represented in the second subspace. Both the first and second spectral components together constitute the complex spectral representation of the signal. The first subspace is the linear span of the first set of basis functions, and the second subspace is the linear span of the second set of basis functions, a part of which is linearly independent of the first set of basis functions.

[0021] In one embodiment, the module that calculates the complex display is a module that calculates the imaginary part of the spectrum of a discrete-time signal based on a real / imaginary conversion, i.e., the real spectrum display of the signal. This conversion is based on exact or approximate mathematical relationships, such as harmonic analysis or equations from heuristic relationships.

[0022] In certain embodiments of the decoder system or encoder system, the first spectral component is obtained by a time-frequency domain conversion of the discrete-time domain signal, preferably by a Fourier transform, such as a discrete cosine transform (DCT), a modified discrete cosine transform (MDCT), a discrete sine transform (DST), a modified discrete sine transform (MDST), a fast Fourier transform (FFT), a prime-factor-based Fourier algorithm, etc. In the first four cases, the second spectral component is obtained by DST, MDST, DCT, and MDCT, respectively. As is well known, the linear span of a cosine periodic over a unit period constitutes a subspace that is not completely included in the linear span of a sine periodic over the same period. Preferably, the first spectral component is obtained by MDCT and the second spectral component is obtained by MDST.

[0023] In one embodiment, the decoder system includes at least one temporal noise shaping module (TNS module, i.e., TNS filter), which is arranged upstream of the upmix stage. Generally speaking, the use of TNS improves the perceived sound quality of signals with transient components, which also applies to embodiments of the decoder system of the present invention having TNS. In conventional L / R and M / S stereo coding, the TNS filter is applied immediately before the inverse transform as the last processing step in the frequency domain. However, in the case of complex prediction stereo coding, it is often advantageous to apply the TNS filter to the downmix signal and the residual signal, i.e., before the upmix matrix. In other words, TNS is applied to the linear combination of the left and right channels, which has several advantages. First, in certain situations, it can be seen that TNS is only advantageous for, for example, the downmix signal. Second, for the residual signal, TNS filtering can be omitted, which means an economical use of the available bandwidth. The TNS filter coefficients only need to be transmitted for the downmix signal. Second, the calculation of the oversampled representation of the downmix signal (for example, to construct the complex frequency domain representation, MDST data is obtained from MDCT data) is required in complex prediction coding, but requires that the time domain representation of the downmix signal can be calculated. This means that the downmix signal can preferably be used as a time sequence of the uniformly obtained MDCT spectrum. When the TNS filter is applied in the decoder after the upmix matrix that converts the downmix / residual representation to the left / right representation, only the sequence of the TNS residual MDCT spectra of the downmix signal is obtained. This makes it very difficult to efficiently calculate the corresponding MDST spectra. This is especially the case when the left / right channels use TNS filters with different characteristics.

[0024] It should be emphasized that whether the time sequence of the MDCT spectrum can be obtained is not an absolute criterion for obtaining an MDST representation that functions as the basis of complex predictive coding. In addition to experimental evidence, this fact can generally be explained by the fact that the TNS is applied only to high frequencies, for example, higher than several kilohertz, so that the residual signal filtered by the TNS approximately corresponds to the unfiltered residual signal of low frequencies. Thus, the present invention can be implemented as a decoder for complex predictive stereo coding in which the TNS filter is arranged upstream of the upmix stage, as described below.

[0025] In one embodiment, the decoder system includes at least one additional TNS module arranged downstream of the upmix stage. By the selector device, the TNS module upstream of the upmix stage or the TNS module downstream of the upmix stage. In certain situations, the calculation of the complex frequency domain representation need not require that the time domain representation of the downmix signal be computable. Further, as described above, the decoder can operate selectively in a direct or simultaneous coding mode without applying complex predictive coding, and it is more suitable to use the TNS module in the conventional location, that is, as one of the last processing steps in the frequency domain.

[0026] In one embodiment, the decoder system is configured to save processing resources and possibly energy by deactivating a module that calculates a second frequency-domain representation of the downmix signal. The downmix signal is partitioned into consecutive time blocks, and each time block is associated with a value of a complex prediction coefficient. This value is determined by a decision for each time block made by an encoder cooperating with the decoder. Further, in this embodiment, the module that calculates the second frequency-domain representation of the downmix signal is configured to deactivate itself if the absolute value of the imaginary part of the complex prediction coefficient is zero or less than a predetermined tolerance value for a given time block. Deactivation of the module means not calculating the second frequency-domain representation of the downmix signal for this time block. If not deactivated, the second frequency-domain representation (e.g., a set of MDST coefficients) is filled with zeros or numbers on the same order as the machine epsilon (rounding unit) of the decoder or other suitable threshold.

[0027] In a further development of the above embodiment, processing resource savings are made at a sub-level of the time blocks into which the downmix signal is partitioned. For example, such a sub-level within a time block is a frequency band, and the encoder determines the value of the complex prediction coefficient for each frequency band within the time block. Similarly, the method of generating the second frequency-domain representation is configured to suppress operations for frequency bands within the time block where the complex prediction coefficient is zero or has a magnitude less than the tolerance value.

[0028] In one embodiment, the first spectral component is a transform coefficient arranged in a time block of transform coefficients, and each block is generated by applying a transform of a time segment of the time-domain signal. Further, the module that calculates the second frequency-domain representation of the downmix signal · obtains a first intermediate component from the first spectral component, · obtains a second intermediate component by constructing a combination of the first spectral components by at least a part of the impulse response. ·configured to obtain a second spectral component from the second intermediate component. By this procedure, as described in detail in U.S. Patent No. 6,980,933B2, particularly in columns 8 to 28 and particularly in formula 41, the second frequency domain representation can be calculated directly from the first frequency domain representation. As those skilled in the art will appreciate, for example, the calculation is not performed in the time domain, as opposed to an inverse transform followed by different transforms.

[0029] In the case of an embodiment of the complex prediction stereo coding according to the present invention, it is speculated that the computational complexity increases only slightly compared to conventional L / R or M / S stereo (significantly less than the increase caused by complex prediction stereo coding in the QMF domain). In this type of embodiment, including the exact calculation of the second spectral component, a delay that is only a few percent longer than that caused by the QMF-based embodiment occurs (assuming a time block length of 1024 samples and comparing with the 961-sample delay of the QMF analysis / synthesis filter bank).

[0030] Preferably, in at least some of the above embodiments, the impulse response is adapted to the transform by which the first frequency domain representation is obtained, more precisely, determined by its frequency response characteristics.

[0031] In some embodiments, the first frequency domain representation of the downmix signal is obtained by a transform applied to one or more analysis window functions (or cutoff functions, such as a rectangular window, a sine window, a Kaiser-Bessel window, etc.), one purpose of which is to achieve temporal segmentation without generating dangerous noise levels or imparting undesirable changes to the spectrum. In some cases, such window functions partially overlap. Next, preferably, the frequency response characteristics of the transform depend on the characteristics of the one or more analysis window functions.

[0032] Referring further to embodiments characterized by the calculation of a second frequency domain representation in the frequency domain, the computational load can be reduced by using an approximate second frequency domain representation. Such approximation can be achieved by not requiring completeness of the information on which the calculation is based. For example, according to the teachings of U.S. Patent No. 6,980,933 B2, the first frequency domain data from three time blocks, namely the block simultaneous with the output block, the preceding block, and the subsequent block, is required for the exact calculation of the second frequency domain representation of the downmix signal in a block. For the purpose of complex predictive coding according to the present invention, by omitting or replacing with zero the data from the subsequent block and / or the preceding block (due to the operation of the module, i.e., not contributing to the delay), a suitable approximation can be obtained such that the calculation of the second frequency domain representation is based on only one or two time blocks. As a note, the omission of the input data means a rescaling of the second frequency domain representation in the sense that, for example, it no longer represents the same power, but as described above, it can be used as the basis for complex predictive coding as long as it is calculated in an equivalent manner on both the encoder side and the decoder side. Indeed, this kind of rescaling is compensated by the corresponding change in the prediction coefficient values.

[0033] Still other approximation methods for calculating spectral components that make up part of the second frequency domain representation of the downmix signal involve combining at least two components from the first frequency domain representation. The latter components are adjacent in time and / or frequency. Alternatively, the latter components can be combined by finite impulse response (FIR) filtering in a relatively small number of steps. For example, in a system using a time block size of 1024, such an FIR filter can include 2, 3, 4, etc. taps. An explanation of this type of approximate calculation method can be found, for example, in U.S. Patent Application Publication No. 2005 / 0197831A1. It is convenient to base the second spectral component of a time block on only a combination of the first spectral components of the same time block by using a window function that gives a relatively small weight in the vicinity of each time block boundary, for example, when using a non-rectangular function, but the same amount of information cannot be obtained for the outermost components. The approximation error that may occur due to such a practice can be reduced to some extent or masked by the shape of the window function.

[0034] In one embodiment of a decoder designed to output a time domain stereo signal, it is possible to switch directly or between simultaneous stereo encoding and complex prediction encoding. This can be achieved by providing: · A switch that can operate selectively as a (signal-unchanging) pass-through stage or as a sum / difference conversion; · An inverse conversion stage that performs a frequency / time conversion; and · A selector device that inputs a directly (or simultaneously) encoded signal or a signal encoded by complex prediction to the inverse conversion stage. As those skilled in the art will notice, since there is such flexibility on the decoder side, the encoder has the freedom to choose between conventional direct or simultaneous encoding and complex predictive encoding. Thus, this embodiment can ensure that at least the same audio quality level as that of conventional direct L / R stereo encoding or simultaneous M / S stereo encoding is maintained if the audio quality level of the latter cannot be exceeded. Therefore, the decoder according to this embodiment can be regarded as a superset of the related art.

[0035] Another group of embodiments of the decoder system calculates the second spectral component of the second frequency domain representation via the time domain. More precisely, the inverse transform of the transform for which the first spectral component has been obtained (or can be obtained) is applied, and then a different transform with the second spectral component as the output is performed. Specifically, MDST is performed after inverse MDCT. In such an embodiment, in order to reduce the number of transforms and inverse transforms, the output of the inverse MDCT is sent to both MDST and the output terminal of the decoding system (optionally preceded by yet another processing step).

[0036] In the case of an embodiment of complex predictive stereo encoding according to the present invention, it is speculated that the computational complexity increases only slightly compared to conventional L / R or M / S stereo (significantly less than the increase caused by complex predictive stereo encoding in the QMF domain).

[0037] As a further development of the embodiment mentioned in the above paragraph, the upmix stage may have a further inverse transform stage for processing the side signal. Then, the time domain representation of the side signal generated by the further inverse transform stage and the time domain representation of the downmix signal generated by the aforementioned inverse transform are supplied to the sum / difference stage. Again, from the perspective of computational complexity, conveniently, the latter signal is supplied to both the above-mentioned sum / difference stage and a different transform stage.

[0038] In one embodiment, a decoder designed to output a time-domain stereo signal is capable of switching between direct L / R stereo encoding or simultaneous M / S stereo encoding and complex prediction stereo encoding. This can be achieved by comprising the following. That is, · A switch that can operate as a pass-through stage or as a sum / difference stage; · A further inverse transform stage for calculating the time-domain representation of the side signal; · The inverse transform stage is connected to a further sum / difference stage that is upstream of the upmix and downstream of the switch (preferably when the switch is activated and functions as a pass filter, as in the case of decoding a stereo signal generated by complex prediction encoding), or to a combination of the downmix signal from the switch and the side signal from the weighted adder (preferably when the switch is activated and functions as a sum / difference stage, as in the case of decoding a directly encoded stereo signal), via a selector device. As those skilled in the art will appreciate, this gives the encoder the freedom to select between conventional direct or simultaneous encoding and complex prediction encoding, i.e., it can guarantee a sound quality level that is at least equal to that of direct or simultaneous stereo encoding.

[0039] In one embodiment, an encoder system according to a second aspect of the present invention has an estimator that estimates complex prediction coefficients for the purpose of reducing or minimizing the signal power or average signal power of a residual signal. The minimization is preferably performed over a certain time period, preferably over an encoding time segment or time block or time frame. The square of the amplitude can be used as a measure of the instantaneous signal power, and the integral of the square of the amplitude over one hour interval can be used as a measure of the average signal power over that time interval. Preferably, the complex prediction coefficients can be determined for each time block and for each frequency band. That is, the value is set to reduce the average power (i.e., total energy) of the residual signal in that time block and frequency band. Specifically, a module that estimates parametric stereo encoding parameters such as IID, ICC, and IPD or similar parameters provides an output that can calculate complex prediction coefficients according to mathematical relationships known to those skilled in the art.

[0040] In one embodiment, the encoding stage of the encoder system further functions as a pass-through stage to enable direct stereo encoding. In situations where direct stereo encoding is expected to provide higher audio quality, by selecting this, the encoder system can ensure that the encoded stereo signal has at least the same audio quality as direct encoding. Similarly, in situations where a large computational load caused by complex prediction encoding is not desirable even if the audio quality is significantly improved, the encoder system has an option to save computational resources readily available. The decision between direct real prediction encoding and complex prediction encoding in the coder generally is based on the principle of rate / distortion optimization.

[0041] In one embodiment, the encoder system has a module that calculates a second frequency domain representation directly based on the first spectral component (i.e., without applying an inverse transform in the time domain and without using the time domain data of the signal). Regarding the corresponding embodiment of the decoder system described above, this module has a similar configuration, i.e., it has similar but different-order processing operations and is configured such that the encoder outputs data suitable for the decoder-side input. For the purpose of explaining this embodiment, it is assumed that the stereo signal to be encoded has mid and side channels or is converted to this configuration, and the encoding stage is configured to receive a first frequency domain representation. The encoding stage has a module that calculates a second frequency domain representation of the mid channel. (The first and second frequency domain representations referred to here are as defined above; specifically, the first frequency domain representation may be an MDCT representation, and the second frequency domain representation may be an MDST representation.) The encoding stage further has a weighted adder that calculates a residual signal as a linear combination weighted by the real and imaginary parts of the complex prediction coefficients, composed of the side signal and the two frequency domain representations of the mid signal. The mid signal, or preferably its first frequency domain representation, is directly used as the downmix signal. In this embodiment, further, an estimator determines the values of the complex prediction coefficients for the purpose of minimizing the power of the residual signal or the average signal power. The final operation (optimization) is performed by feedback control, and if necessary, further by feedback control in which the estimator receives the residual signal obtained by the current prediction coefficient value to be adjusted, or feedforwardly, directly on the left / right channels of the original stereo signal or by calculations performed on the mid / side channels. A feedforward method in which the complex prediction coefficients are calculated directly (in particular, non-iteratively or non-feedbackly) based on the first and second frequency domain representations of the mid signal and the first frequency domain representation of the side signal is preferred. As a note, after determining the complex prediction coefficients, a decision is made directly, considering the quality obtained in each option (preferably, for example, the perceptual quality considering the signal-to-mask effect), whether to perform direct, simultaneous real prediction coding or complex prediction coding.Therefore, the above statement should not be construed as meaning that there is no feedback mechanism in the encoder.

[0042] In one embodiment, the encoder system has a module that calculates a second frequency-domain representation of the mid (i.e., downmix) signal via the time domain. The implementation details regarding this embodiment are the same, at least as far as the calculation of the second frequency-domain representation is concerned, and can be carried out in the same manner as the corresponding decoder embodiment. In this embodiment, the encoding stage has the following: · A sum-difference stage that converts the stereo signal into a mid channel and a side channel; · A conversion stage that provides a frequency-domain representation of the side channel and a complex-valued (i.e., oversampled) frequency-domain representation of the mid channel; and · A weighted adder that calculates a residual signal using complex prediction coefficients as weights. Here, the estimator receives the residual signal and determines complex prediction coefficients that reduce or minimize the power or average power of the residual signal, possibly in a feedback control format. However, preferably, the estimator receives the stereo signal to be encoded and determines the prediction coefficients based thereon. Using the critically sampled frequency-domain representation of the side channel is advantageous from the perspective of computational economy. This is because in this embodiment, the side channel is not multiplied by a complex number. Preferably, the conversion stage may include an MDCT stage and an MDST stage configured in parallel. Both have the time-domain representation of the mid channel as an input. In this way, an oversampled frequency-domain representation of the mid channel and a critically sampled frequency-domain representation of the side channel are generated.

[0043] Note that the methods and apparatuses disclosed in this section can be applied to the encoding of signals having more than two channels with appropriate modifications within the capabilities of those skilled in the art, including ordinary experiments. Such modifications to multi-channel operability can be made, for example, in accordance with Sections 4 and 5 of the paper by J. Herre et al. cited above.

[0044] In yet another embodiment, the features of the above two or more embodiments can be combined, unless they are clearly non-complementary. Just because two features are described in different claims does not mean they cannot be combined. Similarly, in yet another embodiment, features that are not necessary or essential for the desired purpose may be omitted. As an example, the decoding system according to the present invention may be implemented without an inverse quantization stage if the encoded signal to be processed is not quantized or is already in a form suitable for processing in the upmix stage.

Brief Description of the Drawings

[0045] The present invention will be further described by the embodiments described in the following section with reference to the accompanying drawings.

Figure 1A

Figure 1B

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14A

Figure 14B

Figure 14C

Figure 15

Figure 16

Embodiments for Carrying Out the Invention

[0046] I. Decoder System FIG. 2 shows a decoding system for decoding a bit stream having at least one complex prediction coefficient value α = α R + iα I . The MDCT representation of the stereo signal has a downmix M channel and a residual D channel. The real part of the prediction coefficient and α R and the imaginary part α I is quantized and / or jointly coded. However, preferably, the real and imaginary parts are quantized independently and uniformly, generally with a step size of 0.1 (dimensionless number). According to the MPEG standard, the resolution of the frequency band used for the complex prediction coefficients does not have to be the same as the resolution of the scale factor band (sfb, i.e., a group of MDCT lines using the same MDCT quantization step size and quantization range). In particular, the frequency band resolution of the prediction coefficients is acoustically psychologically appropriate, such as the Bark scale. The demultiplexer 201 is configured to extract these MDCT representations and prediction coefficients (part of the illustrated control information) from the supplied bitstream. In fact, the bitstream encodes control information above the complex prediction coefficients, such as instructions on whether to decode it in the predictive or non-predictive mode and TNS information. The TNS information contains the values of the TNS parameters used by the TNS (synthesis) filter of the decoder system. When the same set of TNS parameters is used for multiple TNS filters, such as both channels, it is more economical to receive this information in the form of bits indicating the identity of the set of parameters than to receive two sets of parameters separately. For example, information on whether to apply TNS before or after the upmix stage may also be included based on two optional psychoacoustic evaluations. Furthermore, the control information indicates the individually limited bandwidths of the downmix signal and the residual signal. For each channel, the frequency band above the bandwidth limit is not decoded and is set to zero. In some cases, the energy content of the highest frequency band is small, so it is already zero when quantized. In normal practice (refer to the max_sfb parameter of the MPEG standard), the same bandwidth limit must be used for both the downmix signal and the residual signal. However, the residual signal has significantly more energy content confined to the low frequency band than the downmix signal. Therefore, by imposing a dedicated bandwidth upper limit on the residual signal, it is possible to reduce the bit rate without significantly degrading the audio quality.For example, this is adjusted by two independent max_sfb parameters for the downmix signal and the residual signal, encoded in the bitstream.

[0047] In this embodiment, the MDCT representation of the stereo signal is segmented into consecutive time frames (i.e., time blocks) that include a fixed number of data points (e.g., 1024 points), one of a plurality of fixed numbers of data points (e.g., 128 points or 1024 points), or a variable number of points. As is known to those skilled in the art, the MDCT is critically sampled. The output of the decoding system, shown in the right portion of the figure, is a time-domain stereo signal having left L and right R channels. The inverse quantization module 202 is configured to process the bitstream input to the decoding system into two bitstreams corresponding to, respectively, the downmix channel and the residual channel obtained after demultiplexing the original bitstream as necessary. The inverse quantized channel signals are provided to a switching assembly 203 that can operate in a pass-through mode corresponding to a

Number

Number

[0048] Assuming that the switching assembly 203 is in the pass - through mode here, in this embodiment, the inverse - quantized channel signals pass through their respective TNS filters 204. The TNS filters 204 are not essential for the operation of the decoding system and can also be replaced by pass - through elements. After that, the signals are supplied to a second switching assembly 205 having the same function as the switching assembly 203 arranged upstream. As described above, when the input signal is input and set to the pass - through mode, the output of the second switching assembly 205 is a down - mixed channel signal and a residual channel signal. The down - mixed signal is represented by a temporally continuous MDCT spectrum, but is supplied to a real - imaginary conversion 206 configured to calculate the MDST spectrum of the down - mixed signal. In this embodiment, one MDST frame is based on three MDCT frames, one previous frame, one current (i.e., simultaneous) frame, and one subsequent frame. It is symbolically shown that the input side of the real - imaginary conversion 206 has a delay component (Z -1 ,Z).

[0049] The MDST representation of the down - mixed signal obtained from the real - imaginary conversion 206 is weighted by the imaginary part α I of the prediction coefficient, and the real part α R is added to the MDCT representation of the downmix signal weighted by the MDCT representation of the residual signal. The two additions and multiplications are performed by adders and multipliers that (functionally) constitute the weight adders 210, 211. These are supplied with the values of the complex prediction coefficients α that were encoded in the bitstream first received by the decoder system. One complex prediction coefficient is determined per time frame. The complex prediction coefficients may be determined more frequently, or one may be determined per frequency band in the frame. The frequency bands are partitions motivated acoustically psychologically. As will be described later with respect to the encoding system of the present invention, the complex prediction coefficients may not need to be determined so frequently. The real / imaginary conversion 206 is synchronized with the weight adder so that the current MDST frame of the downmix channel signal is combined with the respective simultaneous MDCT frames of the downmix channel signal and the residual channel signal. The sum of these three signals is the side signal S = Re{αM}+D. In this equation, M includes both the MDCT representation and the MDST representation of the downmix signal, i.e., M = M MDCT -iM MDST is. D = DMDCT is a real value. In this way, a stereo signal having a downmix channel and a side channel is obtained, and the sum-difference conversion 207 obtains from this stereo signal

Number

[0050] A possible implementation of the real / imaginary conversion 206 is described in detail in the applicant's U.S. Patent No. 6,980,933 B2 as described above. According to Equation 41 described in the above document, the conversion can be represented as a finite impulse response filter. For example, for an even number of points,

Number

Number

[0051] It is possible to further reduce the amount of input data for the calculation. For the sake of explanation, in the figure, the real-to-imaginary conversion 206 and its upstream connection, which are the parts indicated by "A", may be replaced by a simplified variant. Two of them, A' and A'', are shown in FIG. 10. Variant A' gives an approximation of the imaginary representation of the signal. Here, the MDST calculation only considers the current frame and the previous frame. Referring to the above formula in this paragraph, for p = 0,..., N - 1, X III (p) is set to 0 (index III indicates a later time frame). Since variant A' does not require the MDCT spectrum of the later frame as input, the MDST calculation does not introduce a time delay. Obviously, this approximation causes a certain reduction in the accuracy of the obtained MDST signal, but it also suggests that the energy of this signal decreases. As a property of predictive coding, the latter can be completely compensated by increasing α I .

[0052] Variant A'' is shown in FIG. 10. This uses only the MDCT data of the current time frame as input. The MDST representation obtained by variant A'' is less accurate than that obtained by variant A'. On the other hand, variant A'' operates with zero delay like variant A' and has low computational complexity. As described above, as long as the same approximation is used in the encoder system and the decoder system, there is no impact on the waveform coding characteristics.

[0053] As a note, regardless of which of variant A, A' or A'', or further developed versions of these are used, only the part where the imaginary part of the complex prediction coefficient of the MDST spectrum is not zero, that is, α I ≠ 0 needs to be calculated. In a practical situation, this means that the absolute value of the imaginary part of the coefficient, |α I It can be understood to mean that | is greater than a predetermined threshold value. This predetermined threshold value relates to the unit round-off of the hardware used. If the imaginary parts of the coefficients in all frequency bands within a time frame are zero, there is no need to calculate the MDST data for that frame. Therefore, again, the real / imaginary conversion 206 is configured not to generate an MDST output, so that when the value of |α I is very small, a response can be made. This can save computing resources. However, in an embodiment that generates one frame of MDST data using frames above the current frame, when the next time frame related to non-zero prediction coefficients occurs, the unit upstream of the conversion 206 must continue to operate even if the MDST spectrum is not required, so that there is sufficient input data for the real / imaginary conversion 206. In particular, the second switching assembly 205 must continue to transfer the MDCT spectrum. This is of course the next time block.

[0054] Returning to FIG. 2, assume that both switching assemblies 203 and 205 are each set to the pass-through mode, and the function of the decoding system will be described. As described herein, the decoder system can decode signals that are not prediction-encoded. For this use, the second switching assembly 205 is set to the sum-and-difference mode, and as shown in the figure, the selector device 208 is set to the lower position so that the signal is directly input to the inverse transform 209 from the source point between the TNS filter 204 and the second switching assembly 205. For proper decoding, the signal has the L / R format at the appropriate source point. Therefore, in order to always supply the correct mid (i.e., downmix) signal to the real / imaginary transform (e.g., not more concisely by the left signal), it is preferable to set the second switching assembly 205 to the sum-and-difference mode when decoding a non-prediction-encoded stereo signal. As described above, prediction encoding is replaced by conventional direct encoding or simultaneous encoding of multiple frames, for example, based on data rate vs. audio quality determination. The result of such a decision is sent from the encoder to the decoder in various ways, for example, by the value of dedicated indicator bits in each frame or by the presence or absence of prediction coefficient values. To prove these facts, the role of the first switching assembly 203 can be easily realized. In fact, in the non-prediction-encoding mode, the decoder system can process both signals by direct (L / R) stereo encoding and signals by simultaneous (M / S) encoding. By operating the first switching assembly 203 in either the pass-through mode or the sum-and-difference mode, it is possible to always provide a source point together with the directly encoded signal. Obviously, the switching assembly 203 converts an input signal in the M / S format to an output signal in the L / R format (supplied to an optional TNS filter 204) when functioning in the sum-and-difference stage.

[0055] The decoder system receives a signal indicating whether to decode a certain time frame in a predictive coding mode or a non-predictive coding mode by the decoder system. The non-predictive mode is signaled by the value of dedicated indicator bits in each frame or by the presence (or value being zero) of prediction coefficients. The predictive mode can be signaled in a similar manner. A particularly advantageous embodiment enables fallback without overhead, but uses the reserved fourth value of the 2-bit field ms_mask_present (see MPEG-2 AAC, ISO / IEC 13818-7 document). This is transmitted for each time frame and is defined as follows, as follows:

Table 1

[0056] Figure 4 shows a decoder system of a general configuration, which is similar to that shown in Figure 2 but includes at least two different configurations. First, the system of Figure 4 includes switches 404 and 411 that enable the application of processing steps including frequency domain modification upstream and / or downstream of the upmix stage. This is realized, on the one hand, by a first set of frequency domain modifiers 403 (depicted as a TNS synthesis filter in this figure) provided together with a first switch 404, which is downstream of the inverse quantization module 401 and the first switching assembly 402 and upstream of a second switching assembly 405 arranged immediately upstream of the upmix stages 406, 407, 408, 409. On the other hand, the decoder system includes a second set of frequency domain modifiers 410 provided together with a second switch 411, which is downstream of the upmix stages 406, 407, 408, 409 and upstream of the inverse transform stage 412. Advantageously, as shown in the figure, each frequency domain modifier is arranged in parallel with a pass-through line connected to the input side of the frequency domain modifier upstream and to the associated switch downstream. With this configuration, signal data is always supplied to the frequency domain modifier, enabling processing in the frequency domain based not only on the current time frame but also on more time frames. The decision on whether to apply the first set of frequency domain modifiers 403 or the second set of frequency domain modifiers 410 may be made by the encoder (sent in the bitstream), or based on whether predictive coding is applied, or based on other criteria suitable for the actual situation. As an example, when the frequency domain modifier is a TNS filter, the first set 403 is advantageous for the use with certain types of signals, while the second set 410 is advantageous for the use with other types of signals. If the result of this selection is encoded in the bitstream, the decoder system appropriately activates each set of the TNS filters.

[0057] To facilitate understanding of the decoder system shown in FIG. 4, it is explicitly noted that the decoding of the direct (L / R) coded signal is α = 0 (suggesting that the pseudo L / R is the same as the L / R and that the side channel and the residual channel are not different), the first switching assembly 402 is in the pass mode, the second switching assembly is in the sum / difference mode, and it is performed when the signal is in the M / S format between the second switching assembly 405 in the upmixing stage and the sum / difference stage 409. At this time, since the upmixing stage is an effective pass step, it is not important whether the first set of frequency domain modifiers or the second set of frequency domain modifiers is activated (using each switch 404, 411).

[0058] FIG. 3 shows a decoder system according to an embodiment of the present invention, which represents different approaches to the supply of MDST data required for upmixing, in relation to the decoder systems of FIGS. 2 and 4. Similar to the decoder system already described, the system of FIG. 3 has an inverse quantization module 301, a first switching assembly 302 operable in a pass-through mode or a sum / difference mode, and a TNS (synthesis) filter 303. These are all arranged in series from the input end of the decoder system. The modules downstream of this point are selectively utilized by two second switches 305, 310. These second switches preferably operate simultaneously such that both are in the upper position or the lower position as shown. At the output end of the decoder system, there is a sum / difference stage 312, and immediately upstream of it, there are two inverse MDCT modules 306, 311 that convert the MDCT domain representation of each channel into a time domain representation.

[0059] In complex prediction decoding, a decoder system is supplied with a downmix / residual stereo signal and a bitstream encoding complex prediction coefficients. The first switching assembly 302 is set to the pass-through mode, and the second switches 305, 310 are set to the upper positions. Downstream of the TNS filter, different processing is performed on the two channels of the (inverse quantized, TNS filtered MDCT) stereo signal. The downmix channel, on the one hand, is supplied to a multiplier and adder 308. The multiplier and adder 308 adds the MDCT representation of the downmix channel weighted by the real part α R of the prediction coefficient to the MDCT representation of the residual channel. On the other hand, it is supplied to one of a plurality of MDCT conversion modules 306. The time-domain representation of the downmix channel M is the output from the inverse MDCT conversion module 306 and is supplied to both the final sum / difference stage 312 and the MDST conversion module 307. Thus, using the time-domain representation of the downmix channel doubly is advantageous from the viewpoint of computational complexity. The MDST representation of the downmix channel obtained in this way is further supplied to another multiplier and adder 309. This multiplier and adder 309 weights by the imaginary part α I of the prediction coefficient and then adds this signal to the linear combination output from the adder 308. Therefore, the output of the adder 309 is the side-channel signal S = Re{αM}+D. Similarly, the multipliers and adders 308, 309 are coupled to the decoder system shown in FIG. 2 to form a weighted multi-signal adder that takes as inputs the MDCT and MDST representations of the downmix signal, the MDCT representation of the residual signal, and the complex prediction coefficient values. In this embodiment, downstream of this point, only the path through the inverse MDCT conversion module 311 remains before the side-channel signal is supplied to the final sum / difference stage 312.

[0060] The synchronization required in the decoder system can be achieved by making the transform length and window shape applied in both inverse MDCT conversion modules 306 and 311 the same. This is already in practical use in frequency-selective M / S and L / R coding. Combining an embodiment of the inverse MDCT module 306 with an embodiment of the MDST module 307 results in a one-frame delay. Therefore, five optional delay blocks 313 (or software instructions that achieve this effect in the case of computer implementation) are provided, and the part of the system on the right side of the dashed line can be delayed by one frame with respect to the part on the left side as needed. Obviously, delay blocks are provided at all intersections between the dashed line and the connection lines, but the connection between the inverse MDCT module 306 and the MDST conversion module 307 is an exception, where a delay that requires compensation occurs.

[0061] Calculating the MDST data for one time frame requires data from one frame of the time-domain display. However, the inverse MDCT conversion is based on one frame (the current frame), two consecutive frames (preferably the previous frame and the current frame), or three consecutive frames (preferably the previous frame, the current frame, and the subsequent frame). Due to the well-known time-domain alias cancellation (TDAC) related to MDCT, the three-frame option achieves a complete overlap of the input frames and is the most (and in some cases completely) accurate, at least for frames containing time-domain aliases. Obviously, the three-frame inverse MDCT operates with a one-frame delay. By allowing the use of an approximate time-domain display as the input to the MDST conversion, this delay can be avoided, thereby avoiding the need to compensate for the delay between different parts of the decoder system. In the two-frame option, overlap / add enabling TDAC is performed in the first half of the frame, and aliases only exist in the second half. In the one-frame option, since there is no TDAC, aliases occur throughout the frame. However, the MDST display realized in this way and used as a daytime signal with complex predictive coding can provide sufficient quality.

[0062] The decoding system shown in FIG. 3 can operate in two non-predictive decoding modes. To directly decode the L / R encoded stereo signal, the second switches 305, 310 are set to the lower position and the first switching assembly 302 is set to the pass-through mode. Thus, this signal is in L / R format upstream of the sum / difference stage 304. The sum / difference stage 304 converts this to M / S format. In this M / S format, the inverse MDCT transform and the final sum / difference operation are performed. To decode the stereo signal provided in the simultaneous M / S encoded format, the first switching assembly 302 is set to the sum / difference mode, and the signal is made to be in L / R format between the first switching assembly 302 and the sum / difference stage 304. The L / R format is more suitable than the M / S format from the viewpoint of TNS filtering. The processing downstream of the sum / difference stage 304 is the same as in the case of direct L / R decoding.

[0063] FIG. 14 (14A to 14C) are three block diagrams showing a decoder according to an embodiment of the present invention. Unlike other block diagrams attached to this application, the connection lines in FIG. 14 indicate multi-channel signals. Specifically, such connection lines are configured to transmit stereo signals having combinations of left / right, mid / side, downmix / residual, pseudo left / pseudo right channels and others.

[0064] Figure 14A shows a decoder system that decodes the frequency domain representation of an input signal (shown as an MDCT representation for the purposes of this figure). The decoder system is configured to supply, as its output, a time domain representation of a stereo signal. This representation is generated based on the input signal. To enable decoding of an input signal encoded by complex predictive stereo coding, the decoder system is provided with an upmix stage 1410. However, it is also possible to process an input signal encoded in another format and, in some cases, switching between multiple coding formats over time, such as an input signal where a sequence of time frames encoded by complex predictive coding is followed by a time portion encoded by direct left / right coding. The functionality of the decoder system for processing different coding formats is realized by providing a connection line (pass-through) in parallel with the upmix stage 1410. By means of a switch 1411, it is possible to select whether to supply to a decoder module arranged further downstream either the output from the upmix stage 1410 (lower switch position in the figure) or the unprocessed signal obtained by the connection line (upper switch position in the figure). In this embodiment, an inverse MDCT module 1412 is arranged downstream of the switch. The MDCT module 1412 converts the MDCT representation of the signal into a time domain representation. As an example, the signal supplied to the upmix stage 1410 may be a stereo signal in downmix / residual form. Next, the upmix stage 1410 is applied to determine the side signal and perform sum / difference operations to output a left / right stereo signal (in the MDCT domain).

[0065] FIG. 14B shows a decoder system similar to that shown in FIG. 14A. This system is configured to receive a bitstream as an input signal. The bitstream is first processed by a combined demultiplexer and inverse quantization module 1420. This combined demultiplexer and inverse quantization module 1420 provides, as a first output signal, an MDCT representation of a multichannel stereo signal for further processing, depending on the position of a switch 1422 that performs a function similar to that of the switch 1411 shown in FIG. 14A. More precisely, the switch 1422 determines whether to process the first output from the demultiplexer and inverse quantization by an upmix stage 1421 and an inverse MDCT module 1423 (lower position), or by the inverse MDCT module 1423 only (upper position). The combined demultiplexer and inverse quantization module 1420 also outputs control information. In this case, the control information related to the stereo signal includes data indicating whether the upper or lower position of the switch 1422 is suitable for decoding the signal, or more abstractly, which coding format to decode the stereo signal into. The control information also includes parameters that adjust the characteristics of the upmix stage, such as the value of the complex prediction coefficient α used in complex prediction coding, as already described.

[0066] FIG. 14C has first and second frequency domain modification devices 1431, 1435 disposed upstream and downstream of the upmix stage 1433, respectively, in addition to the same entities as shown in FIG. 14B. For the purposes of this drawing, each frequency domain modification device is illustrated by a TNS filter. However, the term frequency domain modification device can also be understood as a process applicable before and after the upmix stage other than TNS filtering. Examples of frequency domain modification include prediction, noise addition, bandwidth expansion, and non-linear processing. In some cases, for psychoacoustic considerations and similar reasons, including the characteristics of the signal to be processed and / or the settings of such frequency domain modification devices, it is advantageous to apply the frequency domain modification upstream rather than downstream of the upmix stage 1433. In other cases, for similar considerations, the downstream position of the frequency domain modification is preferred over the upstream position. By switches 1432, 1436, the frequency domain modification devices 1431, 1435 are selectively activated according to control information so that the decoder system can select the desired configuration. As an example, FIG. 14C shows a configuration in which the stereo signal from the combined demultiplexer and inverse quantization module 1430 is first processed by the first frequency domain modification device 1431, then supplied to the upmix stage 1433, and finally transferred directly to the inverse MDCT module 1437 without passing through the second frequency domain adjustment device 1435. As described in the Summary of the Invention section, this configuration is preferred over the option of performing TNS after upmixing in complex predictive coding.

[0067] II. ENCODER SYSTEM The encoder system according to the present invention will be described with reference to FIG. 5. FIG. 5 is a block diagram showing an encoder system that encodes a left / right (L / R) stereo signal as an output bit stream by complex predictive coding. This encoder system receives a time-domain or frequency-domain representation of a signal and supplies it to both a downmix stage and a prediction coefficient estimator. The real and imaginary parts of the prediction coefficients are supplied to the downmix stage to control the downmix of the left and right channels and the conversion to the residual channels. Next, the downmix and residual channels are supplied to a final multiplexer MUX. If the signal is not supplied to the encoder as a frequency-domain representation, it is converted to such a representation in the downmix stage or the multiplexer.

[0068] One of the principles of predictive coding is to convert the left / right signal into a mid / side format, that is

Number

Number

[0069] The real part of the prediction coefficient and α R and the imaginary part α I It is quantized and / or jointly coded. However, preferably, the real and imaginary parts are quantized independently and uniformly, generally with a step size of 0.1 (dimensionless number). According to the MPEG standard, the resolution of the frequency band used for the complex prediction coefficients does not have to be the same as that of the scale factor band (sfb, i.e., a group of MDCT lines using the same MDCT quantization step size and quantization range). In particular, the frequency band resolution of the prediction coefficients is acoustically psychologically appropriate, such as the Bark scale. Note that when the transform length changes, the frequency band resolution changes.

[0070] As described above, the encoder system according to the present invention has the freedom to apply predictive stereo coding or not. In the latter case, a fallback to L / R or M / S coding is suggested. Such a decision can be made on a time frame or finer basis, or on a frequency band basis within a time frame. As described above, the negative result of that decision is sent to the decoding entity in various ways, for example, by the value of a dedicated indicator bit in each frame or by the presence or absence (or zero value) of the prediction coefficient values. The positive decision is sent in the same way. A particularly advantageous embodiment enables fallback without overhead, but uses the reserved fourth value of the 2-bit field ms_mask_present (see MPEG-2 AAC, ISO / IEC 131818-7 document). This is transmitted for each time frame and is defined as follows:

Table 2

[0071] The substantial decision may be based on the data rate vs. audio quality principle. (This is common in the case of available MDCT-based audio encoders, but) As a measure of audio quality, data obtained using the psychoacoustic model included in the encoder may also be used. Specifically, some encoder embodiments perform rate-distortion optimization selection of prediction coefficients. Thus, in such embodiments, if the increase in prediction gain does not save enough bits for encoding the residual signal and cannot justify the use of bits required for encoding the prediction coefficients, the imaginary part of the prediction coefficients, and in some cases also the real part, are set to zero.

[0072] Encoder embodiments encode TNS-related information into the bitstream. Such information includes the values of the TNS parameters used by the TNS (synthesis) filter on the decoder side. When using the same set of TNS parameters for both channels, it is more economical to include signaling bits indicating that the parameters are the same than to transmit the two sets of parameters separately. For example, based on two optional psychoacoustic evaluations, information on whether to apply TNS before or after the upmix stage may also be included.

[0073] As yet another optional feature, which is potentially beneficial from the point of view of complexity and bit rate, the encoder is configured to use an individually limited bandwidth for encoding the residual signal. Frequency bands above this limit are not transmitted to the decoder and are set to zero. In some cases, the energy content of the highest frequency band is small and is already zero when quantized. In normal practice (see the max_sfb parameter of the MPEG standard), the use of the same bandwidth limitation is required for both the downmix signal and the residual signal. Here, the inventors have empirically found that the residual signal has significantly more energy content confined to the low frequency band than the downmix signal. Therefore, by imposing a dedicated bandwidth upper limit on the residual signal, it is possible to reduce the bit rate without significantly degrading the audio quality. For example, this is achieved by transmitting with two independent max_sfb parameters, one for the downmix signal and one for the residual signal.

[0074] It should be noted that the problem of optimal decisions such as prediction coefficients, quantization and its encoding, fallback to M / S or L / R mode, TNS filtering, and bandwidth upper limit has been described with reference to the decoder system shown in FIG. 5, but the same applies equally to the embodiments disclosed in the embodiments described with reference to the subsequent drawings.

[0075] Figure 6 shows another encoder system according to the present invention configured to perform complex predictive stereo coding. This system receives as input a time-domain representation of a stereo signal, divided into successive, possibly overlapping time frames and including left and right channels. The sum / difference stage 601 converts this signal into a mid-channel and a side-channel. The mid-channel is supplied to both the MDCT module 602 and the MDST module 603, and the side-channel is supplied only to the MDCT module 604. The predictive coefficient droplet 605 estimates the values of the complex predictive coefficients for each time frame and, optionally, for individual frequency bands within the frame. The value α of the coefficient is supplied as a weight to the weighted adders 606, 607. The weighted adders 606, 607 form the residual signal D as a linear combination of the MDCT and MDST representations of the mid-signal and the MDCT representation of the side-signal. Preferably, the complex predictive coefficients are supplied to the weighted adders 606, 607 represented by the same quantization scheme used when they are encoded into the bitstream. This clearly provides a more faithful reconstruction since both the encoder and the decoder use the same values of the predictive coefficients. The residual signal, the mid-signal (more appropriately called the downmix signal when it appears in combination with the residual signal), and the predictive coefficients are supplied to the combined quantization and multiplexer stage 608. The combined quantization and multiplexer stage 608 encodes these signals and, optionally, further information as an output bitstream.

[0076] Figure 7 is a modified example of the encoder shown in Figure 6. As can be seen from the same symbols in the figure, the configuration of the encoder shown in Figure 7 is the same, but a function of operating directly in the L / R coding fallback mode is added. The encoder system is activated between the complex predictive coding mode and the fallback mode by a switch 710 provided immediately upstream of the combined quantization and multiplexer stage 709. When the switch 710 is in the upper position, the encoder operates in the fallback mode. The mid-side signal is supplied from a point immediately downstream of the MDCT modules 702, 704 to the sum / difference stage 705. The sum / difference stage 705 converts the signal into a left / right signal and then sends it to the switch 710. The switch 710 connects the signal to the combined quantization and multiplexer stage 709.

[0077] Figure 8 is a diagram showing an encoder system according to the present invention. Different from the encoder systems shown in Figures 6 and 7, in this embodiment, the MDST data required for complex predictive coding is obtained directly from the MDCT data, that is, by performing a real / imaginary conversion in the frequency domain. The real / imaginary conversion applies any of the approaches described with respect to the decoder systems of Figures 2 and 4. To perform faithful decoding, it is important to match the calculation method of the decoder with that of the encoder. The same real / imaginary conversion method is used on the encoder side and the decoder side. For the decoder embodiment, the portion A with the real / imaginary conversion 804, surrounded by a dashed line, can be replaced by a similar modified example or by reducing the input time frames used. Similarly, the encoding can be simplified using any of the above approximate approaches.

[0078] At a high level, the encoder system of FIG. 8 has a configuration different from the configuration that would be obtained by simply replacing the MDST module of FIG. 7 with a real / virtual module (appropriately connected). This architecture is clean, robust and computationally economical, and can implement a switching function between predictive coding and direct L / R coding. The input stereo signal is input to the MDCT conversion module 801, and the MDCT conversion module 801 outputs the frequency domain representation of each channel. This is sent to both the final switch 808 that activates the encoder system between the predictive coding mode and the direct coding mode, and the sum / difference stage 802. In direct L / R coding, or in simultaneous M / S coding performed in a time frame where the prediction coefficient α is set to zero, this embodiment only performs MDCT conversion, quantization, and multiplexing on the input signal. The latter two steps are performed by the combined quantization and multiplexer stage 807 arranged at the output end of the system, and a bit stream is supplied. In predictive coding, each channel is further processed between the sum / difference stage 802 and the switch 808. The real / virtual conversion 804 obtains MDST data from the MDCT representation of the mid signal and sends it to both the prediction coefficient estimator 803 and the weighted adder 806. Similar to the encoder systems shown in FIGS. 6 and 7, another weighted adder 805 is used to combine the side signal with the weighted MDCT and MDST representations of the mid signal to form the residual channel signal. The residual channel signal is encoded by the combined quantization and multiplexer stage 807 together with the mid (i.e., downmix) channel signal and the prediction coefficient.

[0079] Referring now to FIG. 9, it will be explained that each embodiment of the encoder system can be combined with one or more TNS (analysis) filters. As described above, it is often advantageous to apply TNS filtering to a signal in downmix format. Thus, as shown in FIG. 9, adapting the encoder system of FIG. 7 to include TNS is done by adding a TNS filter 911 immediately upstream of the combined quantization and multiplexer stage 909.

[0080] Instead of the right / residual TNS filter 911b, two TNS filters (not shown) configured to process the right channel or the residual channel may be provided immediately upstream of the switch 910. In this way, each of the two TNS filters is always supplied with each channel signal, and TNS filtering based on time frames more than just the current frame is possible. As described above, the TNS filter is an example of a frequency domain correction device, and in particular, it is a device based on the processing of more frames than the current time frame. This gains benefits from such an arrangement as much as or more than the TNS filter.

[0081] As another alternative to the embodiment shown in FIG. 9, a TNS filter for selective activation can be configured at one or more points for each channel. This is similar to the configuration of the decoder system shown in FIG. 4, where different sets of TNS filters can be connected by switches. Thereby, for each time frame, the most suitable stage for TNS filtering can be selected. In particular, with respect to switching between the complex prediction stereo coding mode and other coding modes, it is advantageous to switch between different TNS locations.

[0082] FIG. 11 shows a modification based on the encoder system of FIG. 8, which obtains a second frequency domain representation of the downmix signal by means of a real / imaginary conversion 1105. Similar to the decoder system shown in FIG. 4, this encoder system also includes a frequency domain modifier module that can be selectively activated, one of which 1102 is provided upstream of the downmix stage and one 1109 is provided downstream thereof. The frequency domain modules 1102, 1109 are illustrated by TNS filters in this figure, but can be connected to each signal path using four switches 1103a, 1103b, 1109a and 1109b.

[0083] III. Non-device embodiments Embodiments of the third and fourth aspects of the present invention are shown in FIGS. 15 and 16. FIG. 15 shows a method of decoding a bitstream into a stereo signal and has the following steps: 1. Input a bitstream. 2. Inverse quantize the bitstream to obtain first frequency domain representations of the downmix channel and the residual channel of the stereo signal. 3. Calculate a second frequency domain representation of the downmix channel. 4. Calculate a side channel signal based on the three frequency domain representations of the channels. 5. Calculate the stereo signal, preferably in left / right format, based on the side channel and the downmix channel. 6. Output the stereo signal thus obtained. Steps 3 to 5 may be considered as the upmixing process. Each of steps 1 to 6 is similar to the corresponding function of any decoder system disclosed in the previous part of this document, and details regarding implementation can be read from the same part.

[0084] Figure 16 shows a method for encoding a stereo signal into a bitstream signal and has the following steps: 1. Input a stereo signal. 2. Convert the stereo signal into a first frequency domain representation. 3. Determine complex prediction coefficients. 4. Downmix the frequency domain representation. 5. Encode the downmix channel and the residual channel as a bitstream together with the complex prediction coefficients. 6. Output the bitstream. Each of steps 1 to 5 is similar to the corresponding function of any encoder system disclosed in the previous part of this document, and details regarding implementation can be read from the same part.

[0085] Both methods can be expressed as computer-readable instructions in the form of software programs and can be executed by a computer. The scope of protection of the present invention extends to such software and to computer program products for distributing such software.

[0086] IV. Experimental Evaluation The embodiments disclosed herein were experimentally evaluated. The most important parts of the experimental data obtained in this process are summarized below.

[0087] The embodiments used in the experiment had the following characteristics: (i) Each MDST spectrum (in the time frame) was calculated by two-dimensional finite impulse response filtering from the current, previous, and next MDCT spectra. (ii) An auditory psychological model from the USAC stereo encoder was used. (iii) Instead of the PS parameters ICC, CLD, and IPD, the real and imaginary parts of the complex prediction coefficient α were transmitted. The real and imaginary parts were processed separately, limited to the range [-3.0, 3.0], and quantized using a step size of 0.1. They were time-differentially encoded and finally Huffman-coded using the USAC scale factor codebook. The prediction coefficients were updated every one scale factor band, and the frequency resolution became similar to that of MPEG Surround (see, for example, ISO / IEC 23003-1). With this quantization and coding scheme, in a typical configuration with a target bitrate of 96 kb / s, the average bitrate of the stereo side information became approximately 2 kb / s. (iv) Since there are only three possible values for the 2-bit ms_mask_present bitstream element, the bitstream format was modified without breaking the current USAC bitstream. By using a fourth value indicating complex prediction, a fallback mode of basic mid / side coding was allowed without wasting bits (see the previous subsection of this disclosure for this).

[0088] A listening test was conducted using the MUSHRA method with 8 test items having a sampling rate of 48 kHz and played back on headphones. Each test had 3, 5, or 6 subjects participating.

[0089] The impact due to different MDST approximations was evaluated, showing the actual complexity vs. audio quality trade-off between these options. The results are shown in FIGS. 12 and 13. The former shows the absolute scores obtained, and the latter shows the differential scores with respect to 96s USAC cplf, i.e., for MDCT domain unified stereo coding by complex prediction using the current MDCT frame to calculate the MDST approximation. It can be seen that the audio quality gain achieved by MDCT-based unified stereo coding increases when using a computationally more complex approach to calculate the MDST spectrum. Considering the average over the entire test, the single-frame based system 96s USAC cplf significantly increases the coding efficiency with respect to conventional stereo coding. Similarly, in the case of 96s USAC cp3f, i.e., for MDCT domain unified stereo coding by complex prediction using the current, previous, and next MDCT frames to calculate the MDST, even better results are obtained.

[0090] V. Embodiments Furthermore, the present invention can be implemented as follows.

[0091] A decoder system for decoding a bitstream signal into a stereo signal by complex prediction stereo coding, comprising: An inverse quantization stage (202, 401) for providing a first frequency domain representation of a downmix signal (M) and a residual signal (D) based on the bitstream, each frequency domain representation having a first spectral component representing the spectral content of the corresponding signal represented in a first subspace of a multi-dimensional space, the first spectral component being transform coefficients arranged in a time frame of transform coefficients, each block being generated by applying a transform to a time segment of a time domain signal, the inverse quantization stage; and Configured to generate the stereo signal based on the downmix signal and the residual signal, arranged downstream of the inverse quantization stage: A module (206; 408) that calculates a second frequency domain representation of the downmix signal based on the first frequency domain representation of the downmix signal, wherein the second frequency domain representation has a second spectral component representing spectral content of the signal represented in a second subspace of the multi-dimensional space that includes a part of the multi-dimensional space not included in the first subspace, and the module is configured to obtain a first intermediate component from the first spectral component; obtain a second intermediate component by constructing a combination of the first spectral components with at least a part of an impulse response; and obtain the second spectral component from the second intermediate component; A weighted adder (210, 211; 406, 407) that calculates a side signal based on the first and second frequency domain representations of the downmix signal encoded in the bitstream signal, the first frequency domain representation of the residual signal, and complex prediction coefficients (α); and An upmix stage (206, 207, 210, 211, 406, 407, 408, 409) having a sum / difference stage (207; 409) that calculates the stereo signal based on the first frequency domain representations of the downmix signal and the side signal.

[0092] Furthermore, the present invention can be implemented as follows. That is, a decoder system that decodes a bitstream signal into a stereo signal by complex prediction stereo coding: An inverse quantization stage (301) that provides first frequency domain representations of a downmix signal (M) and a residual signal (D) based on the bitstream signal, wherein each of the first frequency domain representations has a first spectral component representing spectral content of the corresponding signal represented in a first subspace of a multi-dimensional space; and configured to generate the stereo signal based on the downmix signal and the residual signal, disposed downstream of the inverse quantization stage; A module (306, 307) for calculating a second frequency domain representation of the downmix signal based on a first frequency domain representation of the downmix signal, wherein the second frequency domain representation has spectral content of the signal represented in a second subspace of the multi-dimensional space that includes a portion of the multi-dimensional space not included in the first subspace, an inverse transformation stage (306) for calculating a time domain representation of the downmix signal based on the first frequency domain representation of the downmix signal in the first subspace of the multi-dimensional space; and a transformation stage (307) for calculating a second frequency domain representation of the downmix signal based on the time domain representation of the signal. A weighted adder (308, 309) for calculating a side signal based on the first and second frequency domain representations of the downmix signal encoded in the bitstream signal, the first frequency domain representation of the residual signal, and the complex prediction coefficient (α); and An upmix stage (306, 307, 308, 309, 312) having a sum / difference stage (312) for calculating the stereo signal based on the first frequency domain representations of the downmix signal and the side signal.

[0093] Also, the present invention can be implemented as follows. A decoder system having the features described in the claims of an independent decoder system, wherein the module for calculating a second frequency domain representation of the downmix signal: An inverse transformation stage (306) for calculating a time domain representation of the downmix signal and / or the chroma signal based on the first frequency domain representation of each signal in the first subspace of the multi-dimensional space; and A transformation stage (307) for calculating a second frequency domain representation of each signal based on the time domain representation of the signal. Preferably, the inverse transformation stage (306) performs an inverse modified discrete cosine transform, and the transformation stage performs a modified discrete cosine transform.

[0094] In the above decoder system, the stereo signal may be represented in the time domain, and the decoder system may further have the following: (a) A bypass stage used for simultaneous stereo encoding; or (b) A switching assembly (302) arranged between the inverse quantization stage and the upmix stage that can function as either a sum / difference stage used for direct stereo encoding; A further inverse transformation stage (311) arranged in the upmix stage for calculating the time-domain display of the side signal; (a) A further sum / difference stage (304) connected to a point downstream of the switching assembly (302) and upstream of the upmix stage; or (b) A selector device (305, 310) arranged upstream of the inverse transformation stages (306, 301) configured to be selectively connected to either the downmix signal obtained from the switching assembly (302) or the side signal obtained from the weighting adder (308, 309).

[0095] VI. Conclusion Further embodiments of the present invention will be apparent to those skilled in the art upon reading the above description. Although this specification and the drawings disclose embodiments and examples, the present invention is not limited to these specific examples. Numerous modifications and variations can be made without departing from the scope of the present invention as defined in the appended claims.

[0096] Note that the methods and apparatuses disclosed in this application can be appropriately modified within the capabilities of those skilled in the art, including ordinary experiments, and applied to the encoding of signals having more than two channels. It should be emphasized that the signals, parameters, and matrices described in relation to the described embodiments may be frequency-variable or frequency-invariant and / or time-variable or time-invariant. The described calculation steps can be performed for each frequency or all frequencies at once, and all entities can be implemented to have frequency-selective operations. For the purposes of the application, any quantization scheme can be adapted by an auditory psychology model. Further, note that all various sum / difference transforms, namely the conversion from the downmix / residual form to the pseudo L / R form and the L / R-to-M / S conversion and the M / S-to-L / R conversion, are all in the following form

Number

[0097] The systems and methods disclosed herein can be implemented as software, firmware, hardware, or combinations thereof. Some or all components can be implemented as software executed by a digital signal processor or microprocessor, or as hardware or application specific integrated circuits. Such software can be distributed on a computer-readable medium. Computer-readable media include computer storage media and communication media. As is well known to those skilled in the art, computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and other data. Computer storage media includes RAM, ROM, EEPROM, flash memory and other memory technologies, CD-ROM, digital versatile disk (DVD) and other optical disk storage media, magnetic cassettes, magnetic tape, magnetic disk storage and other magnetic storage devices, or any other medium that can be used to store desired information, but is not limited thereto. Further, as is known to those skilled in the art, communication media generally embodies data in a modulated data signal such as computer-readable instructions, data structures, program modules, other carrier waves or other transmission mechanisms, and includes any information delivery media. Note the following appendices. (Appendix 1) A decoder system that provides a stereo signal by complex predictive stereo encoding, An upmix stage configured to generate the stereo signal based on a first frequency domain representation of a downmix signal and a residual signal, each first frequency domain representation having a first spectral component representing spectral content of a corresponding signal represented in a first subspace of a multi-dimensional space, the upmix stage comprising: A module configured to calculate a second frequency domain representation of the downmix signal based on the first frequency domain representation of the downmix signal, the second frequency domain representation having a second spectral component representing spectral content of a signal represented in a second subspace of the multi-dimensional space, the second subspace including a part of the multi-dimensional space not included in the first subspace; A weighted adder configured to calculate a side signal based on the first and second frequency domain representations of the downmix signal encoded in the bitstream signal, the first frequency domain representation of the residual signal, and complex prediction coefficients; An upmix stage having a sum / difference stage configured to calculate the stereo signal based on the first frequency domain representations of the downmix signal and the side signal; The upmix stage further being a decoder system operable in a pass-through mode in which the downmix signal and the residual signal are directly supplied to the sum / difference stage. (Appendix 2) The downmix signal and the residual signal are segmented into time frames, and the upmix stage is configured to receive, for each time frame, a 2-bit data field associated with that frame and operate in an active mode or a pass-through mode according to the value of the data field. The decoder system according to Appendix 1. (Appendix 3) The downmix signal and the residual signal are segmented into time frames. The upmixing stage further receives, in the MPEG bitstream, for each time frame, an ms_mask_present field associated with that frame and is configured to operate in an active mode or a pass-through mode according to the value of the ms_mask_present field. The decoder system according to Appendix 1. (Appendix 4) The upmixing stage further includes an inverse quantization stage disposed upstream of the upmixing stage that provides the first frequency domain representations of the downmix signal and the residual signal based on a bitstream signal. The decoding system according to any one of Appendices 1 to 3. (Appendix 5) The first spectral component has a real value represented in the first subspace. The second spectral component has an imaginary value represented in the second subspace. Optionally, the first spectral component is obtained by one of a discrete cosine transform (DCT) or a modified discrete cosine transform (MDCT). Optionally, the second spectral component is obtained by one of a discrete sine transform (DST) or a modified discrete sine transform (MDST). The decoder system according to any one of Appendices 1 to 4. (Appendix 6) At least one temporal noise shaping (TNS) module disposed upstream of the upmixing stage and at least one additional TNS module disposed downstream of the upmixing stage. And a selector device for selectively activating either (a) the TNS module upstream of the upmixing stage or (b) the additional TNS module downstream of the upmixing stage. (Appendix 7) The downmix signal is partitioned into consecutive time frames, and each time frame is associated with a value of a complex prediction coefficient. The decoder system according to any one of Appendices 1 to 5. (Appendix 7) The downmix signal is partitioned into consecutive time frames, and each time frame is associated with a value of a complex prediction coefficient. The module that calculates the second frequency domain representation of the downmix signal is configured to deactivate itself and not generate an output for that time frame in response to the absolute value of the imaginary part of the complex prediction coefficient being less than a predetermined tolerance value for the time frame. The decoder system according to Supplementary Note 5. (Supplementary Note 8) The downmix signal time frame is further partitioned into frequency bands, and each frequency band is accompanied by a value of the complex prediction coefficient. The module that calculates the second frequency domain representation of the downmix signal is configured to deactivate itself and not generate an output for that frequency band in response to the absolute value of the imaginary part of the complex prediction coefficient being less than a predetermined tolerance value for the frequency band of the time frame. The decoder system according to Supplementary Note 7. (Supplementary Note 9) The first spectral component is a conversion coefficient arranged in a time frame of conversion coefficients, and each block is generated by applying a conversion to a time segment of the time domain signal. The module that calculates the second frequency domain representation of the downmix signal obtains a first intermediate component from the first spectral component, forms a combination of the first spectral components by at least a part of the impulse response to obtain a second intermediate component, and obtains a second spectral component from the second intermediate component. The decoder system according to any one of Supplementary Notes 1 to 8. (Supplementary Note 10) A part of the impulse response is based on the frequency response characteristics of the conversion. Optionally, the frequency response characteristics of the conversion depend on the characteristics of the analysis window function applied to the conversion of the signal to the time segment. The decoder system according to Supplementary Note 9. (Supplementary Note 11) The module that calculates the second frequency domain representation of the downmix signal (a) The simultaneous time frame of the first spectral component (b) The simultaneous and previous time frames of the first spectral component, and (c) configured to determine each time frame of the second spectral component based on one of the simultaneous, previous, and subsequent time frames of the first spectral component The decoder system according to appendix 9 or 10. (Appendix 12) The module for calculating the second frequency domain representation of the downmix signal is configured to calculate an approximate second spectral representation having an approximate second spectral component determined by a combination of at least two spectrally adjacent and / or temporally adjacent first spectral components. The decoder system according to any one of appendices 1 to 11. (Appendix 13) The stereo signal is represented in the time domain, and the decoder system further (a) A switching assembly arranged between the inverse quantization stage and the upmix stage, which can function as either (a) a pass-through stage or (b) a sum / difference stage, thereby switching directly and simultaneously between the encoded stereo input signals, An inverse transform stage configured to calculate a time domain representation of the stereo signal, An upstream selector device arranged upstream of the inverse transform stage and configured to selectively connect it to either (a) a point downstream of the upmix stage where a stereo signal obtained by complex prediction is supplied to the inverse transform stage, or (b) a point downstream of the switching assembly and upstream of the upmix stage where a stereo signal obtained by direct stereo encoding is supplied to the inverse transform stage. The decoder system according to any one of appendices 1 to 12. (Appendix 14) An encoder system for encoding a stereo signal using complex prediction as a signal having a downmix channel, a residual channel, and complex prediction coefficients, An estimator for estimating complex prediction coefficients, (a) Convert the stereo signal into a frequency-domain representation of a downmix signal and a residual signal having a relationship determined by the values of the complex prediction coefficients, and (b) an encoding stage that operates as a pass-through stage and is operable to directly supply the stereo signal to be encoded to a multiplexer. (Appendix 15) Configured to encode a stereo signal by a complex prediction stereo encoding into a bitstream signal, and further receives the output from the encoding stage and the estimator, and further includes a multiplexer that encodes by the bitstream signal. The encoder system according to Appendix 14. (Appendix 16) The estimator determines the complex prediction coefficients by minimizing the power of the residual signal over time or the average power of the residual signal. The encoder system according to Appendix 14 or 15. (Appendix 17) The stereo signal has a downmix channel and a side channel. The encoding stage is configured to receive a first frequency-domain representation of the stereo signal, and the first frequency-domain representation has a first spectral component representing the spectral content of the corresponding signal represented in a first subspace of a multi-dimensional space. The encoding stage further a module that calculates a second frequency-domain representation of the downmix channel based on the first frequency-domain representation of the downmix signal, where the second frequency-domain representation has a second spectral component representing the spectral content of a signal represented in a second subspace of the multi-dimensional space that includes a part of the multi-dimensional space not included in the first subspace. and a weighted adder that calculates a residual signal based on the first and second frequency-domain representations of the downmix channel, the first frequency-domain representation of the side channel, and the complex prediction coefficients. The estimator receives the downmix channel and the side channel and determines the complex prediction coefficients to minimize the power of the residual signal over a period of time or to minimize the average power of the residual signal. An encoder system according to any one of appendices 14 to 16. (Appendix 18) The encoding stage A sum-difference stage that converts the stereo signal into a simultaneous encoding stereo signal having a downmix channel and a side channel, A conversion stage that provides an oversampled frequency domain representation of the downmix channel and a critically sampled frequency domain representation of the side channel, wherein the oversampled frequency domain representation preferably has complex spectral components, the conversion stage; Based on the oversampled frequency domain representation of the downmix channel, the critically sampled frequency domain representation of the side channel, and the complex prediction coefficients, it has a weighted adder that calculates a residual signal. The estimator receives the residual signal and determines the complex prediction coefficients to minimize the power of the residual signal or to minimize the average power of the residual signal. Preferably, the conversion stage has a modified discrete cosine transform MDCT stage arranged in parallel with a modified discrete sine transform MDST stage that provides the oversampled frequency domain representation of the downmix channel together. An encoder system according to any one of appendices 14 to 16. (Appendix 19) A decoding method for providing a stereo signal by complex prediction stereo encoding, Receiving a first frequency domain representation of a downmix signal and a residual signal, each of the first frequency domain representations having a first spectral component representing the spectral content of the corresponding signal represented in a first subspace of a multi-dimensional space; Receiving a control signal; According to the value of the control signal, (a) A step of upmixing the downmix signal and the residual signal using an upmix stage to obtain the stereo signal, comprising: A sub-step of calculating a second frequency domain representation of the downmix signal based on a first frequency domain representation of the downmix signal, wherein the second frequency domain representation has a second spectral component representing spectral content of a signal represented in a second subspace of the multi-dimensional space, which is a part of the multi-dimensional space not included in the first subspace; A sub-step of calculating a side signal based on the first and second frequency domain representations of the downmix signal encoded in the bitstream signal, the first frequency domain representation of the residual signal, and complex prediction coefficients; A sub-step of calculating the stereo signal by applying the first frequency domain representations of the downmix signal and the side signal to sum-difference conversion; (b) A step of interrupting the upmixing step. (Appendix 20) The first spectral component has a real value represented in the first subspace; The second spectral component has an imaginary value represented in the second subspace; Optionally, the first spectral component is obtained by one of a discrete cosine transform (DCT) or a modified discrete cosine transform (MDCT); Optionally, the second spectral component is obtained by one of a discrete sine transform (DST) or a modified discrete sine transform (MDST); The decoding method according to Appendix 19. (Appendix 21) The downmix signal is partitioned into consecutive time frames, and each time frame is associated with a value of complex prediction coefficients; The step of calculating the second frequency domain representation of the downmix signal is interrupted in response to the absolute value of the imaginary part of the complex prediction coefficients being smaller than a predetermined tolerance value of the time frame, and no output is generated for that time frame. The decoding method according to Appendix 20. (Appendix 22) The downmix signal time frame is further partitioned into frequency bands, and each frequency band is accompanied by a value of the complex prediction coefficient. The step of calculating the second frequency domain representation of the downmix signal is interrupted in response to the absolute value of the imaginary part of the complex prediction coefficient being smaller than a predetermined tolerance value of the frequency band of the time frame, and no output is generated for that frequency band. The decoding method according to Appendix 21. (Appendix 23) The first spectral component is a transform coefficient arranged in a time frame of transform coefficients, and each block is generated by applying a transform to a time segment of the time domain signal. The step of calculating the second frequency domain representation of the downmix signal includes a sub-step of obtaining a first intermediate component from the first spectral component, a sub-step of constructing a combination of the first spectral components by at least a part of the impulse response to obtain a second intermediate component, and a sub-step of obtaining the second spectral component from the second intermediate component. The decoding method according to Appendix 20. (Appendix 24) A part of the impulse response is based on the frequency response characteristic of the transform. Optionally, the frequency response characteristic of the transform depends on the characteristic of the analysis window function applied to the transform of the time segment of the signal. The decoding method according to Appendix 23. (Appendix 25) The step of calculating the second frequency domain representation uses, as an input, one of (a) the simultaneous time frame of the first spectral components, (b) the simultaneous and previous time frames of the first spectral components, and (c) the simultaneous, previous, and subsequent time frames of the first spectral components to obtain each time frame of the second spectral components. The decoding method according to Appendix 24. (Appendix 26) The step of calculating the second frequency domain representation of the downmix signal includes the step of calculating an approximate second spectral representation having approximate second spectral components determined by a combination of at least two spectrally and / or temporally adjacent first spectral components. The decoding method according to any one of Appendices 19 to 25. (Appendix 27) The stereo signal is represented in the time domain, and the method further includes omitting the upmixing step according to whether the bitstream signal is encoded by direct stereo encoding or simultaneous stereo encoding, and inverting the bitstream signal to obtain the stereo signal. The decoding method according to any one of Appendices 19 to 26. (Appendix 28) Omitting the step of transmitting the time domain representation of the downmix signal and the step of calculating the side signal according to whether the bitstream is encoded by direct stereo encoding or simultaneous stereo encoding, and further including inverting the frequency domain representation of each channel encoded by the bitstream signal to obtain the stereo signal. The decoding method according to Appendix 27. (Appendix 29) An encoding method for encoding a stereo signal by a bitstream by complex prediction stereo encoding, the method including determining complex prediction coefficients, and converting the stereo signal into a first frequency domain representation of a downmix signal and a residual signal having a relationship determined by the complex prediction coefficients, the first frequency domain representation having first spectral components representing spectral contents of corresponding signals represented in a first subspace of a multi-dimensional space, encoding the downmix channel, the residual channel, and the complex prediction coefficients as the bitstream. (Appendix 30) The step of determining the complex prediction coefficients is performed to minimize the power of the residual signal or the average power of the residual signal over a certain period of time. The encoding method described in Appendix 29. (Appendix 31) The step of defining or recognizing the partition of the stereo signal into time frames, and For each time segment, further comprising the step of encoding or selecting to encode the stereo signal in this time segment by at least one of the options of direct stereo encoding, simultaneous stereo encoding, and complex prediction stereo encoding. When direct stereo encoding is selected, the stereo signal is converted to a frequency domain representation of the left and right channels and encoded as the bitstream. When simultaneous stereo encoding is selected, the stereo signal is converted to a frequency domain representation of the downmix channel and the side channel and encoded as the bitstream. The encoding method described in Appendix 29 or 30. (Appendix 32) The option that provides the highest sound quality according to a predetermined psychoacoustic model is selected. The encoding method described in Appendix 31. (Appendix 33) Further comprising the step of defining or recognizing the partition of the stereo signal into time frames, The stereo signal has a downmix channel and a side channel, The step of converting the stereo signal to a first frequency domain representation of the downmix channel and the residual channel is A sub-step of calculating a second frequency domain representation of the downmix signal based on the first frequency domain representation of the downmix channel, wherein the second frequency domain representation has a second spectral component representing the spectral content of a signal represented in a second subspace of the multi-dimensional space, which is a part of the multi-dimensional space not included in the first subspace. Steps of constructing a residual signal based on the first and second frequency region displays of the downmix channel, the first frequency region display of the side channel, and the complex prediction coefficients. The step of determining the complex prediction coefficients is performed for one time frame at a time by minimizing the average power of the residual signal in each time frame. The encoding method according to Appendix 29 or 30. (Appendix 34) Steps of converting the stereo signal into a simultaneous encoding stereo signal having a downmix channel and a side channel. Steps of converting the downmix channel into an oversampled frequency region display preferably having complex spectral components. Steps of converting the side channel into a critically sampled, preferably real-valued frequency region display. Further comprising steps of calculating a residual signal based on the oversampled frequency region display of the downmix channel, the critically sampled frequency region display of the side channel, and the complex prediction coefficients. The determination of the complex prediction coefficients is performed by feedback control regarding the thus calculated residual signal in order to minimize power or average power. The encoding method according to any one of Appendices 29 to 33. (Appendix 35) The conversion of the downmix channel into an oversampled frequency region display is performed by applying MDCT and MDST and concatenating their outputs. The encoding method according to Appendix 34. (Appendix 36) A computer program product having a computer-readable medium storing instructions for executing the method according to any one of Appendices 19 to 35 when executed by a general-purpose computer.< / nd>

Claims

1. An apparatus for outputting a stereo audio signal having a left channel and a right channel, the apparatus comprising: a demultiplexer that receives an audio bitstream and decodes at least one prediction coefficient from the audio bitstream, the audio bitstream being segmented into frames, and the value of the at least one prediction coefficient being changeable for each frame; a decoder configured to generate a downmix signal and a residual signal from the audio bitstream; an upmixer that operates in a prediction mode or a non-prediction mode and is configured to output the left channel and the right channel as the stereo audio signal; when the upmixer operates in the prediction mode, the residual signal represents a difference between a side signal and a prediction of the side signal, and the upmixer generates the left channel and the right channel from a combination of the downmix signal, the residual signal, and the at least one prediction coefficient; when the upmixer operates in the non-prediction mode, the residual signal represents the side signal, and the upmixer generates the left channel based on a sum of the downmix signal and the residual signal, and generates the right channel based on a difference between the downmix signal and the residual signal; Apparatus.

2. The at least one prediction coefficient reduces or minimizes the energy of the residual signal. The apparatus according to claim 1.

3. The apparatus further comprises a noise shaper configured to shape noise associated with the downmix signal, the noise shaper being disposed upstream of the upmixer. The apparatus according to claim 1.

4. The noise shaper is a temporal noise shaper configured to shape the noise over time. The apparatus according to claim 3.

5. When operating in the prediction mode, the upmixer generates the left channel and the right channel using a filter having three taps. The apparatus according to claim 1.

6. The downmix signal includes a mid signal formed by a linear combination of an original left channel and an original right channel. The apparatus according to claim 1.

7. The at least one prediction coefficient is a real-valued coefficient. The apparatus according to claim 1.

8. The at least one prediction coefficient is a complex-valued coefficient. The apparatus according to claim 1.

9. The upmixer combines the side signal with the downmix signal by adding a version of the downmix signal to a version of the side signal to generate the left channel and subtracting the version of the side signal from the version of the downmix signal to generate the right channel. The apparatus according to claim 1.

10. The apparatus according to claim 1, wherein the upmixer is configured to add the residual signal to the side signal when operating in the prediction mode.

11. The downmix signal is divided into frequency bands, the at least one prediction coefficient includes prediction coefficients for each frequency band, the demultiplexer receives the audio bit stream and is configured to decode each of the prediction coefficients from the audio bit stream, The upmixer generates the left channel and the right channel from a combination of the downmix signal, the residual signal, and each of the prediction coefficients when operating in the prediction mode. The apparatus according to claim 1.

12. A method of outputting a stereo audio signal having a left channel and a right channel, the method comprising: receiving an audio bit stream and decoding at least one prediction coefficient from the audio bit stream, wherein the audio bit stream is segmented into frames and the value of the at least one prediction coefficient can vary for each frame; generating a downmix signal and a residual signal from the audio bit stream in a decoder; upmixing in a prediction mode or a non-prediction mode; outputting the left channel and the right channel as the stereo audio signal, wherein when the upmixing is performed in the prediction mode, the residual signal represents the difference between a side signal and a prediction of the side signal, and the upmixing generates the left channel and the right channel from a combination of the downmix signal, the residual signal, and the at least one prediction coefficient. When performing the upmixing while operating in the unpredictable mode, the residual signal represents the side signal, and the upmixing generates the left channel based on the sum of the residual signal passed through the decoder and the downmix signal, and generates the right channel based on the difference between the residual signal and the downmix signal. Method. **Claim 13** A non-transitory computer-readable medium including instructions that, when executed by a processor, perform the method according to claim 12.

Citation Information

Patent Citations

  • signal processing

    JP2005521921A

  • Methods for Improving Performance of Prediction-Based Multi-Channel Reconstruction

    JP2008517337A

  • Enhanced coding and parameterization in multi-channel downmixed object coding

    JP2010507115A

  • A parametric stereo upmix apparatus, a parametric stereo decoder, a parametric stereo downmix apparatus, a parametric stereo encoder

    WO2009141775A1