Coding of multichannel audio content

The method addresses computational inefficiencies in multi-channel audio playback by encoding a downmix for legacy systems, allowing flexible decoding and reduced bit rates through mid signals and stereo techniques.

JP2025163183APending Publication Date: 2025-10-28DOLBY INTERNATIONAL AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025130421
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2014-04-01
Filing Date
2025-08-05
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing methods require complete decoding of all channels of multi-channel audio content for playback on legacy systems with fewer channels, leading to computational inefficiency.

Method used

A method and system for encoding and decoding multi-channel audio content that allows direct decoding of a downmix suitable for legacy playback systems by using mid signals and additional input audio signals, with stereo decoding techniques to generate output channels for the desired speaker configuration.

Benefits of technology

Enables efficient decoding of multi-channel audio content on legacy systems without the need for full channel decoding, offering flexibility in speaker configurations and bit rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025163183000001_ABST
    Figure 2025163183000001_ABST
Patent Text Reader

Abstract

To provide a decoding method and an encoding method for a multichannel audio content that allows for effective decoding of downmix suitable for legacy playback systems.SOLUTION: A decoding method 100 for decoding N channels includes the steps of: in a first decoding module, decoding M input audio signals 122 into M mid signals 126 suitable for playback on a speaker configuration having M channels; and for each channel exceeding M channels out of the N channels, receiving an additional input audio signal corresponding to one of the M mid signals to decode the input audio signal and its corresponding mid signal, and generating a stereo signal containing first and second audio signals suitable for playback through two of the N channels of the speaker configuration.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates generally to encoding multi-channel audio signals, and more particularly to encoders and decoders for encoding and decoding multiple input signals for playback on a speaker configuration having a certain number of channels. [Background technology]

[0002] Multichannel audio content corresponds to a speaker configuration with a certain number of channels. For example, multichannel audio content may correspond to five front channels, four surround channels, four ceiling channels, and a low-frequency effects (LFE) channel. Such a channel configuration may be referred to as a 5 / 4 / 4.1, 9.1+4, or 13.1 configuration. Sometimes, it is desirable to play encoded multichannel audio content on a playback system with a speaker configuration that has fewer channels, or speakers, than the encoded multichannel audio content. Hereinafter, such a playback system is referred to as a legacy playback system. For example, it may be desirable to play encoded 13.1 audio content on a speaker configuration with three front channels, two surround channels, two ceiling channels, and an LFE channel. Such a channel configuration may also be referred to as a 3 / 2 / 2.1, 5.1+2, or 7.1 configuration. Summary of the Invention [Problem to be solved by the invention]

[0003] According to the prior art, a complete decoding of all channels of the original multi-channel audio content and a subsequent downmix to the channel configuration of the legacy playback system would be required. Clearly, such a configuration is computationally inefficient since all channels of the original multi-channel audio content need to be decoded. Therefore, there is a need for a coding scheme that allows direct decoding of a downmix suitable for legacy playback systems. [Brief explanation of the drawings]

[0004] Exemplary embodiments will now be described with reference to the accompanying drawings. [Figure 1] FIG. 1 illustrates a decoding scheme according to an exemplary embodiment. [Figure 2] FIG. 2 is a diagram showing an encoding method corresponding to the decoding method of FIG. [Figure 3] FIG. 2 illustrates a decoder according to an exemplary embodiment. [Figure 4] A diagram illustrating a first configuration of a decoding module based on an exemplary embodiment. [Figure 5] A diagram illustrating a second configuration of a decoding module based on an exemplary embodiment. [Figure 6] FIG. 2 illustrates a decoder according to an exemplary embodiment. [Figure 7] FIG. 2 illustrates a decoder according to an exemplary embodiment. [Figure 8] FIG. 8 illustrates a high-frequency reconstruction component used in the decoder of FIG. 7. [Figure 9] FIG. 1 illustrates an encoder according to an exemplary embodiment. [Figure 10] FIG. 1 illustrates a first configuration of an encoding module according to an exemplary embodiment. [Figure 11]FIG. showing a second configuration of an encoding module according to an exemplary embodiment. All drawings are schematic and generally show only the parts necessary to clarify the present disclosure. On the other hand, other parts may be omitted or only suggested. Unless otherwise specified, like reference numerals refer to like parts in different drawings. **DETAILED DESCRIPTION OF THE INVENTION**

[0005] In view of the above, it is an object to provide an encoding / decoding method for encoding / decoding multi-channel audio content that allows efficient decoding of downmix suitable for a legacy playback system.

[0006] I. Overview - Decoder According to a first aspect, a decoding method, a decoder, and a computer program product for decoding multi-channel audio content are provided.

[0007] According to an exemplary embodiment, a method in a decoder for decoding a plurality of input audio signals for playback in a speaker configuration having N channels, wherein the plurality of input audio signals represent encoded multi-channel audio content corresponding to at least N channels, the method comprising: receiving M input audio signals, where 1 < M ≤ N ≤ 2M; in a first decoding module, decoding the M input audio signals into M mid signals suitable for playback in a speaker configuration having M channels; for each of the N channels that exceeds M channels, receiving an additional input audio signal corresponding to one of the M mid signals, the additional input audio signal being a side signal or a complementary signal that allows reconstruction of the side signal together with the mid signal and a weighting parameter a; decoding, in a stereo decode module, the additional input audio signal and its corresponding mid signal to generate a stereo signal including first and second audio signals suitable for playback on two of the N channels of the speaker configuration; thereby generating N audio signals suitable for playback on the N channels of said speaker configuration. A method is provided.

[0008] The above method is advantageous in that if the audio content is to be played back on a legacy playback system, the decoder does not need to decode all channels of the multi-channel audio content to form a downmix of the complete multi-channel audio content.

[0009] More specifically, a legacy decoder designed to decode audio content corresponding to an M-channel speaker configuration may simply take M input audio signals and decode them into M mid signals suitable for playback on an M-channel speaker configuration. No further downmixing of the audio content is required on the decoder side. In fact, the downmix suitable for the legacy playback speaker configuration is already prepared and encoded on the encoder side and is represented by the M input signals.

[0010] A decoder designed to decode audio content corresponding to more than M channels may receive additional input audio signals and combine these with corresponding ones of the M mid signals by stereo decoding techniques to arrive at output channels corresponding to the desired speaker configuration. The proposed method is therefore advantageous in that it is flexible with respect to the speaker configuration used for reproduction.

[0011] According to an exemplary embodiment, the stereo decode module is operable in at least two configurations depending on the bit rate at which the decoder receives data, and the method may further include receiving an indication as to which of the at least two configurations to use in decoding the additional input audio signal and its corresponding mid signal.

[0012] This is advantageous in that the present decoding method is flexible with respect to the bitrate used by the encoding / decoding system.

[0013] According to an exemplary embodiment, the step of receiving an additional input audio signal comprises: receiving a pair of audio signals corresponding to joint encoding of an additional input audio signal corresponding to a first of the M mid signals and an additional input audio signal corresponding to a second of the M mid signals; decoding the pair of audio signals to generate the additional input audio signals corresponding to first and second ones of the M mid signals, respectively;

[0014] This has the advantage that the additional input audio signals can be efficiently coded pairwise.

[0015] According to an exemplary embodiment, the additional input audio signal is a waveform-encoded signal including spectral data corresponding to frequencies up to a first frequency and the corresponding mid signal is a waveform-encoded signal including spectral data corresponding to frequencies up to frequencies greater than the first frequency, and decoding the additional input audio signal and its corresponding mid signal in accordance with the first configuration of the stereo decoding module comprises: if the additional audio input signal is in the form of a complementary signal, calculating a side signal for frequencies up to the first frequency by multiplying the mid signal by a weighting parameter a and adding the result of the multiplication to the complementary signal; and upmixing the mid signal and the side signal to generate a stereo signal including first and second audio signals, wherein for frequencies below the first frequency, the upmixing comprises performing an inverse sum-difference transform of the mid signal and the side signal, and for frequencies above the first frequency, the upmixing comprises performing a parametric upmix of the mid signal.

[0016] This is advantageous in that the decoding performed by the stereo decoding module allows decoding of a mid signal and a corresponding additional input audio signal, which is waveform-coded to a frequency lower than the corresponding frequency of the mid signal. In this way, this decoding method allows the encoding / decoding system to operate at a reduced bit rate.

[0017] Performing a parametric upmix of a mid signal generally means that for frequencies above the first frequency, the first and second audio signals are parametrically reconstructed based on the mid signal.

[0018] According to an exemplary embodiment, the waveform-encoded mid-signal includes spectral data corresponding to frequencies up to a second frequency, and the method further comprises: Prior to performing a parametric upmix, the method includes extending the mid signal to a frequency range above the second frequency by performing high frequency reconstruction.

[0019] In this way, the present decoding method allows the encoding / decoding system to operate at a reduced bit rate.

[0020] According to an exemplary embodiment, the additional input audio signal and the corresponding mid signal are waveform-encoded signals including spectral data corresponding to frequencies up to a second frequency, and decoding the additional input audio signal and its corresponding mid signal according to the second configuration of the stereo decoding module comprises: if the additional audio input signal is in the form of a complementary signal, calculating a side signal by multiplying the mid signal by the weighting parameter a and adding the result of the multiplication to the complementary signal; and performing an inverse sum-difference transform of the mid signal and the side signal to generate a stereo signal including first and second audio signals.

[0021] This is advantageous in that the decoding performed by the stereo decoding module also allows decoding of the mid signal and a corresponding additional input audio signal, which is waveform-coded up to the same frequency. In this way, this decoding method allows the encoding / decoding system to operate at high bit rates.

[0022] According to an exemplary embodiment, the method further comprises extending the first and second audio signals of the stereo signal to a frequency range above the second frequency by performing high frequency reconstruction, which is advantageous in that it provides further flexibility in terms of bit rate of the encoding / decoding system.

[0023] According to an exemplary embodiment in which the M mid signals are reproduced on a speaker configuration having M channels, the method further comprises: Extending the frequency range of at least one of the M mid signals by performing high-frequency reconstruction based on high-frequency reconstruction parameters associated with the first and second audio signals of the stereo signal that may be generated from at least one of the M mid signals and its corresponding additional audio input signal.

[0024] This is advantageous in that the quality of the high-frequency reconstructed mid signal can be improved.

[0025] According to an exemplary embodiment in which the additional input audio signal is in the form of a side signal, the additional input audio signal and the corresponding mid signal are waveform-coded using modified discrete cosine transforms having different transform sizes. This is advantageous in that it increases the flexibility in choosing the transform size.

[0026] An exemplary embodiment also relates to a computer program product having a computer-readable medium having instructions for performing any of the encoding methods disclosed above. The computer-readable medium may be a non-transitory computer-readable medium.

[0027] An exemplary embodiment also relates to a decoder that decodes a plurality of input audio signals for playback in a speaker configuration having N channels. The plurality of input audio signals represent encoded multi-channel audio content corresponding to at least N channels, and the decoder: a receiving component configured to receive M input audio signals, where 1 < M ≤ N ≤ 2M, the receiving component; a first decoding module configured to decode the M input audio signals into M mid signals suitable for playback in a speaker configuration having M channels; and a stereo encoding module for each of the N channels that exceeds M channels, the stereo encoding module: receives an additional input audio signal corresponding to one of the M mid signals, the additional input audio signal being a side signal or a complementary signal that allows reconstruction of the side signal together with the mid signal and a weighting parameter a; configured to decode the additional input audio signal and its corresponding mid signal to generate a stereo signal comprising first and second audio signals suitable for playback on two of the N channels of the speaker configuration; The decoder is thereby configured to generate N audio signals suitable for reproduction on the N channels of said loudspeaker configuration.

[0028] II. Overview - Encoders According to a second aspect, there is provided an encoding method, an encoder and a computer program product for decoding multi-channel audio content.

[0029] The second aspect may generally have the same features and advantages as the first aspect.

[0030] According to an exemplary embodiment, a method in an encoder for encoding a plurality of input audio signals representing multi-channel audio content corresponding to K channels, comprising: receiving K input audio signals corresponding to channels of a speaker configuration having K channels; generating, from the K input audio signals, M mid signals and KM output audio signals suitable for playback on a speaker configuration having M channels, <M<K≦2Mであり、 2M-K of the mid signals correspond to 2M-K of the input audio signals; The remaining KM mid signals and KM output audio signals are, for each value of K above M, a stereo encoding module for encoding two of the K input audio signals to generate a mid signal and an output audio signal, the output audio signal being a side signal or a complementary signal that allows reconstruction of the side signal together with the mid signal and a weighting parameter a; encoding the M mid signals into M additional output audio channels in a second encoding module; and including the KM output audio signals and the M additional output audio channels in a data stream for transmission to a decoder.

[0031] According to an exemplary embodiment, the stereo encoding module is operable in at least two configurations depending on a desired bit rate of the encoder, and the method may further include including in the data stream an indication as to which of the at least two configurations was used by the stereo encoding module in encoding two of the K input audio signals.

[0032] According to an exemplary embodiment, the method may further include performing stereo encoding of the KM output audio signals pairwise prior to including them in the data stream.

[0033] According to an exemplary embodiment in which the stereo encoding module operates according to a first configuration, encoding two of the K input audio signals to generate a mid signal and an output audio signal comprises: converting the two input audio signals into a first signal that is a mid signal and a second signal that is a side signal; waveform encoding the first and second signals into first and second waveform-encoded signals, respectively, wherein the second signal is waveform-encoded to a first frequency and the first signal is waveform-encoded to a second frequency greater than the first frequency; subjecting the two input audio signals to parametric stereo encoding to extract parametric stereo parameters that allow reconstruction of the two spectral data of the K input audio signals for frequencies above the first frequency; and including the first and second waveform-encoded signals and the parametric stereo parameters in the data stream.

[0034] According to an exemplary embodiment, the method further comprises: for frequencies below the first frequency, converting the waveform-coded second signal, which is a side signal, into a complementary signal by multiplying the waveform-coded first signal, which is a mid signal, by a weighting factor a and subtracting the result of the multiplication from the second waveform-coded signal; and including the weighting parameter a in the data stream.

[0035] According to an exemplary embodiment, the method further comprises: subjecting the first signal, being a mid-signal, to high-frequency reconstruction encoding to generate high-frequency reconstruction parameters that enable high-frequency reconstruction of the first signal above the second frequency; and including the high-frequency reconstruction parameters in the data stream.

[0036] According to an exemplary embodiment in which the stereo encoding module operates according to a second configuration, encoding two of the K input audio signals to generate a mid signal and an output audio signal includes: converting the two input audio signals into a first signal that is a mid signal and a second signal that is a side signal; waveform encoding the first and second signals into first and second waveform-encoded signals, respectively, wherein the first and second signals are waveform-encoded to a second frequency; and including said first and second waveform-encoded signals.

[0037] According to an exemplary embodiment, the method further comprises: converting the waveform-coded second signal, a side signal, into a complementary signal by multiplying the waveform-coded first signal, a mid signal, by a weighting factor a and subtracting the result of the multiplication from the second waveform-coded signal; and including the weighting parameter a in the data stream.

[0038] According to an exemplary embodiment, the method further comprises: subjecting each of the two of the K input audio signals to high-frequency reconstruction encoding to generate high-frequency reconstruction parameters that enable high-frequency reconstruction of the two of the K input audio signals above the second frequency; and including the high-frequency reconstruction parameters in the data stream.

[0039] Exemplary embodiments also relate to a computer program product having a computer-readable medium with instructions for performing the encoding method of the exemplary embodiments. The computer-readable medium may be a non-transitory computer-readable medium.

[0040] Exemplary embodiments also relate to an encoder for encoding a plurality of input audio signals representing multi-channel audio content corresponding to K channels, the encoder comprising: a receiving component configured to receive K input audio signals corresponding to channels of a speaker configuration having K channels; a first encoding module configured to generate, from the K input audio signals, M mid signals and KM output audio signals suitable for playback on a speaker configuration having M channels, the first encoding module comprising: <M<K≦2Mであり、 2M-K of the mid signals correspond to 2M-K of the input audio signals; The first encoding module includes KM stereo encoding modules configured to generate the remaining KM mid signals and KM output audio signals, each stereo encoding module comprising: a first encoding module configured to encode two of the K input audio signals to generate a mid signal and an output audio signal, the output audio signal being a side signal or a complementary signal that allows reconstruction of the side signal together with the mid signal and a weighting parameter a; a second encoding module configured to encode the M mid signals into M additional output audio channels; a multiplexing component configured to include the KM output audio signals and the M additional output audio channels in a data stream for transmission to a decoder. [Example]

[0041] III. ILLUSTRATIVE EMBODIMENTS A stereo signal with left (L) and right (R) channels can be represented differently according to different stereo coding schemes. According to a first coding scheme, referred to in this paper as left-right coding (LR coding), the input channels L, R and output channels A, B of a stereo transform component are related by the following equation: L=A; R=B In other words, LR coding simply implies passing the input channels through. A stereo signal represented by an L and an R channel is said to have an L / R representation or to be in L / R format.

[0042] According to a second coding scheme, referred to herein as sum-difference coding (or mid-side coding, "MS coding"), the input and output channels of a stereo transform component are related by the following equation: A=0.5(L+R); B=0.5(LR) In other words, MS coding involves calculating the sum and difference of the input channels. This is referred to herein as performing a sum-difference transform. For this reason, channel A may be considered the mid signal (sum signal M) of the first and second channels L and R, and channel B may be considered the side signal (difference signal) of the first and second channels L and R. When a stereo signal is subjected to sum-difference coding, the signal is said to have a mid / side (M / S) representation or to be in mid / side (M / S) format.

[0043] From the decoder's point of view, the corresponding expression is L=(A+B); R=(AB) is.

[0044] Converting a stereo signal in mid / side format to L / R format is referred to herein as performing an inverse sum-difference transform.

[0045] The mid-side coding scheme can be generalized to a third coding scheme, referred to in this paper as "enhanced MS coding" (or enhanced sum-difference coding). In enhanced MS coding, the input and output channels of a stereo transform component are related by the following equation: A=0.5(L+R); B=0.5(L(1-a)-R(1+a)) L=(1+a)A+B; R=(1-a)AB where a is a weighting parameter. The weighting parameter may be variable in time and frequency. In this case, signal A may be considered the mid signal, and signal B may be considered the modified side signal or the complementary side signal. In particular, for a=0, the enhanced MS coding scheme reduces to mid-side coding. When a stereo signal is subjected to enhanced mid / side coding, the signal is said to have a mid / complement / a representation (M / c / a) or to be in mid / complement / a format.

[0046] According to the above, the complementary signal may be converted into a side signal by multiplying the corresponding mid signal by a parameter a and adding the result of the multiplication to the complementary signal.

[0047] Figure 1 shows a decoding method 100 in a decoding system according to an exemplary embodiment. A data stream 120 is received by a receiving component 102. The data stream 120 represents encoded multi-channel audio content corresponding to K channels. The receiving component 102 may demultiplex and dequantize the data stream 120 to form M input audio signals 122 and K - M input audio signals 124. Here, it is assumed that M < K.

[0048] The M input audio signals 122 are decoded by a first decoding module 104 into M mid signals 126. The M mid signals are suitable for reproduction in a speaker configuration having M channels. The first decoding module 104 can generally operate according to any known decoding method for decoding audio content corresponding to M channels. Thus, when the decoding system is a legacy or low-computation decoding system that only supports reproduction in a speaker configuration having M channels, the M mid signals can be reproduced on the M channels of the speaker configuration without the need to decode all K channels of the original audio content.

[0049] For a decoding system that supports reproduction in a speaker configuration having N channels where M < N ≤ K, the decoding system may apply the M mid signals 126 and at least a part of the K - M input audio signals 124 to a second decoding module 106. The second decoding module 106 generates N output audio signals 128 suitable for reproduction in a speaker configuration having N channels.

[0050] Each of the KM input audio signals 124 corresponds to one of the M mid signals 126 according to one of two alternatives. According to the first alternative, the input audio signal 124 is a side signal corresponding to one of the M mid signals 126, and the mid signal and the corresponding input signal form a stereo signal expressed in mid / side format. According to the second alternative, the input audio signal 124 is a complementary signal corresponding to one of the M mid signals 126, and the mid signal and the corresponding input signal form a stereo signal expressed in mid / complement / a format. Thus, according to the second alternative, the side signal can be reconstructed from the mid signal and the complementary signal together with the weighting parameter a. When the second alternative is used, the weighting parameter a is included in the data stream 120.

[0051] As described in more detail below, some of the N output audio signals 128 of the second decoding module 106 may correspond directly to some of the M mid signals 126. Additionally, the second decoding module may include one or more stereo decoding modules, each of which operates on the M mid signals 126 and its corresponding input audio signal 124 to generate a pair of output audio signals, each pair of which is suitable for playback on two of the N channels of a speaker configuration.

[0052] 2 shows an encoding scheme 200 of an encoding system corresponding to the decoding scheme 100 of FIG. 1. K input audio signals 228, where K>2, corresponding to channels of a speaker configuration having K channels, are received by a receiving component (not shown). The K input audio signals are input to a first encoding module 206. Based on the K input audio signals 228, the first encoding module 206 generates M mid signals 226 and KM output audio signals 224 suitable for playback on a speaker configuration having M channels, where M <K≦2Mである。

[0053] Generally, as will be explained in more detail below, some of the M mid signals 226, typically 2M-K of the mid signals 226, correspond to respective ones of the K input audio signals 228. In other words, the first encoding module 206 generates some of the M mid signals 226 by passing some of the K input signals 228 through.

[0054] The remaining KM of the M mid signals 226 are generally generated by downmixing, i.e., linearly combining, the input audio signals 228 not passed through the first encoding module 206. In particular, the first encoding module may downmix these input audio signals 228 pairwise. For this purpose, the first encoding module may include one or more (typically KM) stereo encoding modules. Each stereo encoding module operates on a pair of input audio signals 228 to generate a mid signal (i.e., a downmix or sum signal) and a corresponding output audio signal 224. The output audio signal 224 corresponds to the mid signal according to either of the two alternatives discussed above. That is, the output audio signal 224 is either a side signal or a complementary signal that allows reconstruction of the side signal together with the mid signal and a weighting parameter a. In the latter case, the weighting parameter a is included in the data stream 220.

[0055] The M mid signals 226 are then input to the second encoding module 204, where they are encoded into M additional output audio signals 222. The second encoding module 204 may operate according to any known encoding scheme for encoding audio content corresponding to the M channels.

[0056] The N M output audio signals 224 and the M additional output audio signals 222 from the first encoding module are then quantized and included by the multiplexing component 202 in a data stream 220 for transmission to a decoder.

[0057] In the encoding / decoding scheme described with reference to Figures 1-2, a suitable downmix of K-channel audio content to M-channel audio content is performed on the encoder side (by the first encoding module 206), thus achieving efficient decoding of K-channel audio content for playback in a channel configuration with M channels, or more generally N channels, where M < N < K.

[0058] Exemplary embodiments of the decoder are described below with reference to FIGS.

[0059] 3 shows a decoder 300 configured for decoding multiple input audio signals for playback on a speaker configuration having N channels. The decoder 300 includes a receiving component 302, a first decoding module 104, and a second decoding module 106 that includes a stereo decoding module 306. The second decoding module 106 may further include a high frequency extension component 308. The decoder 300 may also include a stereo conversion component 310.

[0060] The operation of the decoder 300 is described below. The receiving component 302 receives a data stream 320, i.e., a bitstream, from the encoder. The receiving component 302 may include, for example, a demultiplexing component that demultiplexes the data stream 320 into its component parts and a dequantizer for dequantizing the received data.

[0061] The received data stream 320 includes a plurality of input audio signals. Generally, the plurality of input audio signals may correspond to encoded multi-channel audio content corresponding to a speaker configuration having K channels, where K≥N.

[0062] In particular, the data stream 320 includes M input audio signals 322, where 1 < M < N. In the illustrated example, M is equal to 7 and there are seven input audio signals 322. However, in other examples, it may be other numbers such as 5. Further, the data stream 320 includes N - M audio signals 323, from which N - M input audio signals 324 can be decoded. In the illustrated example, N is equal to 13 and there are six additional input audio signals 324.

[0063] The data stream 320 may further have an additional audio signal 321, which typically corresponds to an encoded LFE channel.

[0064] According to one example, a pair of the N - M audio signals 323 may correspond to a pair of N - M input audio signals 324 that are jointly encoded. The stereo conversion component 310 may decode such a pair of the N - M audio signals 324 to generate a corresponding pair of the N - M input audio signals 324. For example, the stereo conversion component 310 may perform decoding by applying MS or improved MS decoding to a pair of the N - M audio signals 323.

[0065] The M input audio signals 322 and, if available, the additional audio signal 321 are input to the first decoding module 104. As discussed with reference to FIG. 1 , the first decoding module 104 decodes the M input audio signals 322 into M mid signals 326 suitable for playback on a speaker configuration having M channels. As shown in this example, the M channels may correspond to a center front speaker (C), a left front speaker (L), a right front speaker (R), a left surround speaker (LS), a right surround speaker (RS), a left ceiling speaker (LT), and a right ceiling speaker (RT). The first decoding module 104 further decodes the additional audio signal 321 into an output audio signal 325, which typically corresponds to a low frequency effects (LFE) speaker.

[0066] 1, each of the additional input audio signals 324 corresponds to one of the mid signals 326 in that it is a side signal corresponding to the mid signal or a complementary signal corresponding to the mid signal. By way of example, a first one of the input audio signals 324 may correspond to the mid signal 326 associated with a left front speaker, a second one of the input audio signals 324 may correspond to the mid signal 326 associated with a right front speaker, etc.

[0067] The M mid signals 326 and the N M audio input audio signals 324 are input to a second decoding module 106 which generates N audio signals 328 suitable for playback on an N channel speaker configuration.

[0068] The second decoding module 106 maps the mid signals 326 that do not have corresponding residual signals to corresponding channels of the N-channel speaker configuration, optionally via a high-frequency reconstruction component 308. For example, a mid signal corresponding to a center front speaker (C) of an M-channel speaker configuration may be mapped to the center front speaker (C) of the N-channel speaker configuration. The high-frequency reconstruction component 308 is similar to that described below with reference to Figures 4 and 5.

[0069] The second decoding module 106 includes N M stereo decoding modules 306, one for each pair of a mid signal 326 and a corresponding input audio signal 324. Generally, each stereo decoding module 306 performs joint stereo decoding to generate stereo audio signals that map to two of the channels of an N-channel speaker configuration. As an example, a stereo decoding module 306 that takes as input a mid signal corresponding to a left front speaker (L) of a seven-channel speaker configuration and its corresponding input audio signal 324 generates stereo audio signals that map to the two left front speakers ("Lwide" and "Lscreen") of a 13-channel speaker configuration.

[0070] The stereo decode module 306 can operate in at least two configurations, depending on the data transmission rate (bit rate) at which the encoder / decoder system operates, i.e., the bit rate at which the decoder 300 receives data. A first configuration may correspond to a moderate bit rate, e.g., approximately 32-48 kbps per stereo decode module 306. A second configuration may correspond to a high bit rate, e.g., a bit rate greater than 48 kbps per stereo decode module 306. The decoder 300 receives an indication as to which configuration to use. For example, such an indication may be signaled to the decoder 300 by the encoder via one or more bits in the data stream 320.

[0071] 4 illustrates the stereo decode module 306 when functioning according to a first configuration corresponding to a medium bitrate. The stereo decode module 306 includes a stereo transform component 440, various time / frequency transform components 442, 446, 454, a high frequency reconstruction (HFR) component 448, and a stereo upmix component 452. The stereo decode module 306 is constrained to take as input a mid signal 326 and a corresponding input audio signal 324. The mid signal 326 and the input audio signal 324 are assumed to be represented in the frequency domain, typically the modified discrete cosine transform (MDCT) domain.

[0072] To achieve a moderate bit rate, at least the bandwidth of the input audio signal 324 is limited. More precisely, the input audio signal 324 is a waveform-encoded signal containing spectral data corresponding to frequencies up to a first frequency k1. The mid signal 326 is a waveform-encoded signal containing spectral data corresponding to frequencies up to a frequency greater than the first frequency k1. In some cases, to save additional bits that need to be sent in the data stream 320, the bandwidth of the mid signal 326 is also limited. Thereby, the mid signal 326 contains spectral data up to a second frequency k2 greater than the first frequency k1.

[0073] The stereo conversion component 440 converts the input signals 326, 324 into a mid / side representation. As discussed further above, the mid signal 326 and the corresponding input audio signal 324 may be represented in mid / side format or mid / complement / a format. In the former case, the stereo conversion component 440 passes the input signals 326, 324 without any modification because the input signals are already in mid / side format. In the latter case, the stereo conversion component 440 passes the mid signal 326. Meanwhile, the input audio signal 324, which is a complementary signal, is converted into a side signal for frequencies up to the first frequency k1. More precisely, the stereo conversion component 440 determines the side signal for frequencies up to the first frequency k1 by multiplying the mid signal 326 by a weighting parameter a (received from the data stream 320) and adding the result of the multiplication to the input audio signal 324. As a result, the stereo conversion component thus outputs the mid signal 326 and the corresponding side signal 424.

[0074] In this regard, it is worth noting that if the mid signal 326 and the input audio signal 324 are received in mid / side format, mixing of the signals 324, 326 is not performed in the stereo conversion component 440. As a result, the mid signal 326 and the input audio signal 324 may be encoded by MDCT transforms having different transform sizes. However, if the mid signal 326 and the input audio signal 324 are received in mid / complement / a format, the MDCT encoding of the mid signal 326 and the input audio signal 324 is constrained to the same transform size.

[0075] If the mid signal 326 has limited bandwidth, i.e., if the spectral content of the mid signal 326 is constrained to frequencies up to the second frequency k2, the mid signal 326 is subjected to high frequency reconstruction (HFR) by the high frequency reconstruction component 448. HFR generally refers to a parametric technique that reconstructs the spectral content of a signal for its low frequencies (in this case, frequencies below the second frequency k2) and the high frequencies (in this case, frequencies above the second frequency k2) based on parameters received from the encoder in the data stream 320. Such high frequency reconstruction techniques are known in the art and include, for example, the spectral band replication (SBR) technique. The HFR component 448 thus outputs the mid signal 426 with spectral content up to the maximum frequency represented in the system, where the spectral content above the second frequency k2 is parametrically reconstructed.

[0076] The high frequency reconstruction component 448 typically operates in the quadrature mirror filter (QMF) domain. Therefore, before performing high frequency reconstruction, the mid signal 326 and the corresponding side signal 424 are first transformed into the time domain by a time / frequency transform component 442, which typically performs an inverse MDCT transform, and then transformed into the QMF domain by a time / frequency transform component 446.

[0077] The mid signal 426 and the side signal 424 are then input to a stereo upmix component 452, which generates a stereo signal 428 represented in L / R format. The side signal 424 only has spectral content for frequencies up to a first frequency k1, and the stereo upmix component 452 treats frequencies below and above the first frequency k1 differently.

[0078] More specifically, for frequencies up to the first frequency k1, the stereo upmix component 452 converts the mid signal 426 and the side signal 424 from mid / side format to L / R format, i.e., performs an inverse sum-difference transform for frequencies up to the first frequency k1.

[0079] For frequencies above the first frequency k1 for which no spectral data is provided for the side signal 424, the stereo upmix component 452 parametrically reconstructs the first and second components of the stereo signal 428 from the mid signal 426. Generally, the stereo upmix component 452 receives parameters extracted for this purpose at the encoder side via the data stream 320 and uses these parameters for the reconstruction. Generally, any known technique for parametric stereo reconstruction may be used.

[0080] In view of the above, the stereo signal 428 output by the stereo upmix component 452 thus has spectral content up to the highest frequency represented in the system, where the spectral content above the first frequency k1 is parametrically reconstructed. Like the HFR component 448, the stereo upmix component 452 typically operates in the QMF domain. Thus, the stereo signal 428 is converted to the time domain by a time-to-frequency transform component 454 to generate a time-domain represented stereo signal 328.

[0081] 5 illustrates the stereo decoding module 306 when operating according to a second configuration supporting high bit rates. The stereo decoding module 306 includes a first stereo transform component 540, various time / frequency transform components 542, 546, and 554, a second stereo transform component 452, and high frequency reconstruction (HFR) components 548a and 548b. The stereo decoding module 306 is constrained to take as input a mid signal 326 and a corresponding input audio signal 324. It is assumed that the mid signal 326 and the input audio signal 324 are represented in the frequency domain, typically the modified discrete cosine transform (MDCT) domain.

[0082] For high bit rates, the bandwidth constraints of the input signals 326, 324 are different from those for medium bit rates. More precisely, the mid signal 326 and the input audio signal 324 are waveform-coded signals containing spectral data corresponding to frequencies up to a second frequency k2. In some cases, the second frequency k2 may correspond to the maximum frequency represented by the system. In other cases, the second frequency k2 may be lower than the maximum frequency represented by the system.

[0083] The mid signal 326 and the input audio signal 324 are input to a first stereo conversion component 540 for conversion to a mid / side representation. The first stereo conversion component 540 is similar to the stereo conversion component 440 of FIG. 4. The difference is that if the input audio signal 324 is in the form of a complementary signal, the first stereo conversion component 540 converts the complementary signal to a side signal for frequencies up to a second frequency k2. Thus, the stereo conversion component 540 outputs a mid signal 326 and a corresponding side signal 524, both of which have spectral content up to the second frequency k2.

[0084] The mid signal 326 and the corresponding side signal 524 are then input to a second stereo conversion component 552. The second stereo conversion component 552 forms a sum and a difference of the mid signal 326 and the side signal 524 to convert the mid signal 326 and the side signal 524 from mid / side format to L / R format. In other words, the second stereo conversion component performs an inverse sum-difference transform to generate a stereo signal having a first component 528a and a second component 528b.

[0085] Preferably, the second stereo conversion component 552 operates in the time domain. Therefore, prior to being input to the second stereo conversion component 552, the mid signal 326 and the side signal 524 may be converted from the frequency domain (MDCT domain) to the time domain by the time-to-frequency conversion component 542. Alternatively, the second stereo conversion component 552 may operate in the QMF domain. In such a case, the order of components 546 and 552 in FIG. 5 is reversed. This is advantageous in that the mixing that occurs in the second stereo conversion component 552 does not impose any further constraints on the MDCT transform sizes for the mid signal 326 and the input audio signal 324. Furthermore, as discussed above, if the mid signal 326 and the input audio signal 324 are received in mid / side format, they may be encoded by MDCT transforms using different transform sizes.

[0086] If the second frequency k2 is lower than the highest represented frequency, the first and second components 528a, 528b of the stereo signal may be subjected to high-frequency reconstruction (HFR) by high-frequency reconstruction components 548a, 548b. The high-frequency reconstruction components 548a, 548b are similar to the high-frequency reconstruction component 448 of FIG. 4. However, it is worth noting that in this case, a first set of high-frequency reconstruction parameters is received via data stream 230 and used in the high-frequency reconstruction of the first component 528a of the stereo signal, and a second set of high-frequency reconstruction parameters is received via data stream 230 and used in the high-frequency reconstruction of the second component 528b of the stereo signal. Thus, the high-frequency reconstruction components 548a, 548b output the first and second components 530a, 530b of the stereo signal that contain spectral data up to the highest frequency represented in the system, where the spectral content above the second frequency k2 is parametrically reconstructed.

[0087] Preferably, the high frequency reconstruction is performed in the QMF domain, and therefore, prior to being subjected to the high frequency reconstruction, the first and second components 528a, 528b of the stereo signal may be transformed into the QMF domain by a time / frequency transform component 546.

[0088] The first and second components 530a, 530b of the stereo signal output from the high frequency reconstruction component 548 may then be transformed to the time domain by a time / frequency transform component 554 to generate the stereo signal 328 represented in the time domain.

[0089] Figure 6 shows a decoder 600 configured for decoding multiple input audio signals contained in a data stream 620 for playback in a speaker configuration having 11.1 channels. The structure of the decoder 600 may be generally similar to that shown in Figure 3. The difference is that the speaker configuration shows fewer channels than Figure 3, which shows a speaker configuration having 13.1 channels, and has an LFE speaker, three front speakers (center C, left L, and right R), four surround speakers (left side Lside, left rear Lback, right side Rside, right rear Rback), and four ceiling speakers (upper left front LTF, upper left rear LTB, upper right front RTF, and upper right rear RTB).

[0090] 6, the first decoding component 104 outputs seven mid signals 626, which may correspond to the speaker configurations of channels C, L, R, LS, RS, LT, and RT. In addition, there are four additional input audio signals 624a-d. Each of the additional input audio signals 624a-d corresponds to one of the mid signals 626. By way of example, input audio signal 624a may be a side signal or a complement signal corresponding to an LS mid signal, input audio signal 624b may be a side signal or a complement signal corresponding to an RS mid signal, input audio signal 624c may be a side signal or a complement signal corresponding to an LT mid signal, and input audio signal 624d may be a side signal or a complement signal corresponding to an RT mid signal.

[0091] In the illustrated embodiment, the second decoding module 106 includes four stereo decoding modules 306 of the type shown in Figures 4 and 5. Each stereo decoding module 306 takes as input one of the mid signals 626 and a corresponding additional input audio signal 624a-d and outputs a stereo audio signal 328. For example, based on the LS mid signal and the input audio signal 624a, the second decoding module 106 may output a stereo signal corresponding to the L side and L back speakers. Further examples will be apparent from the figures.

[0092] Additionally, the second decoding module 106 acts as a pass-through for three of the mid signals 626, here corresponding to the C, L, and R channels. Depending on the spectral bandwidth of these signals, the second decoding module 106 may perform high-frequency reconstruction using the high-frequency reconstruction component 308.

[0093] 7 shows how a legacy or low-complexity decoder 700 decodes multi-channel audio content in a data stream 720 corresponding to a speaker configuration with K channels for playback on a speaker configuration with M channels. By way of example, K may be equal to 11 or 13, and M may be equal to 7. The decoder 700 includes a receiving component 702, a first decoding module 704, and a high-frequency reconstruction module 712.

[0094] As further described with reference to data stream 120 in FIG. 1, data stream 720 may generally have M input audio signals 722 (see signals 122 and 322 in FIGS. 1 and 3) and KM additional input audio signals (see signals 124 and 324 in FIGS. 1 and 3). Optionally, data stream 720 may also have an additional audio signal 721, which typically corresponds to the LFE channel. Because decoder 700 supports a speaker configuration with M channels, receiving component 702 only extracts M input audio signals 722 (and additional audio signal 721, if present) from data stream 720 and discards the remaining KM additional input audio signals.

[0095] The M input audio signals 722 and the additional audio signal, exemplified here by seven audio signals, are then input to the first decoding module 104. The first decoding module 104 decodes the M input audio signals 722 into M mid signals 726 corresponding to the channels of an M-channel speaker configuration.

[0096] If the M mid signals 726 only contain spectral content up to a certain frequency below the maximum frequency represented by the system, the M mid signals 726 may be subjected to high frequency reconstruction by the high frequency reconstruction module 712.

[0097] 8 shows an example of such a high frequency reconstruction module 712. The high frequency module 712 has a high frequency reconstruction component 848 and various time / frequency transform components 842, 846, 854.

[0098] The mid signal 726 input to the HFR module 712 is subjected to high frequency reconstruction by an HFR component 848. The high frequency reconstruction is preferably performed in the QMF domain. Thus, the mid signal 726, typically in the form of an MDCT spectrum, may be converted to the time domain by a time / frequency transform component 842 and then converted to the QMF domain by a time / frequency transform component 846 prior to being input to the HFR component 848.

[0099] The HFR component 848 generally operates in the same manner as, for example, the HFR components 448, 548 of Figures 4 and 5, in that it uses the spectral content of the input data for lower frequencies together with parameters received from the data stream 720 to parametrically reconstruct the spectral content for higher frequencies. However, depending on the bit rate of the encoder / decoder system, the HFR component 848 may use different parameters.

[0100] As described with reference to FIG. 5, for the high bit rate case, for each mid signal with a corresponding additional input audio signal, the data stream 720 includes a first set of HRF parameters and a second set of HRF parameters (see the description of items 548a and 548b in FIG. 5). Although the decoder 700 does not use the additional input audio signal corresponding to the mid signal, the HFR component 848 may use a combination of the first and second sets of HRF parameters when performing high-frequency reconstruction of the mid signal. For example, the high-frequency reconstruction component 848 may use a downmix, such as an average or linear combination of the first and second sets of HRF parameters.

[0101] In this manner, the HFR component 854 outputs a mid signal 828 with extended spectral content. The mid signal 828 may then be converted to the time domain by a time-to-frequency conversion component 854 to provide an output signal 728 with a time-domain representation.

[0102] Exemplary embodiments of the encoder are described below with reference to FIGS.

[0103] Figure 9 shows an encoder 900 that conforms to the general structure of Figure 2. The encoder 900 includes a receiving component (not shown), a first encoding module 206, a second encoding module 204, and a quantization and multiplexing component 902. The first encoding module 206 may further include a high frequency reconstruction (HFR) encoding component 908 and a stereo encoding module 906. The decoder 900 may further include a stereo conversion component 910.

[0104] The operation of the encoder 900 will now be described. The receiving component receives K input audio signals 928 corresponding to the channels of a K-channel speaker configuration. For example, the K channels may correspond to the channels of a 13-channel configuration as described above. Additionally, an additional channel 925, typically corresponding to an LFE channel, may be received. The K channels are input to a first encoding module 206, which generates M mid signals 926 and KM output audio signals 924.

[0105] The first encoding module 206 includes KM stereo encoding modules 906. Each of the KM stereo encoding modules 906 takes as input two of the K input audio signals and produces one of the mid signals 926 and one of the output audio signals 924, as will be described in more detail below.

[0106] The first encoding module 206 further maps the remaining input audio signals that are not input to one of the stereo encoding modules 906 to one of M mid signals 926, optionally via an HFR encoding component 908. The HFR encoding component 908 is similar to that described with reference to Figures 10 and 11.

[0107] The M mid signals 926, optionally together with an additional input audio signal 925, typically representing the LFE channel, are input to a second encoding module 204 as described above with reference to Figure 2, for encoding into M output audio channels 922.

[0108] Before being included in the data stream 920, the KM output audio signals 924 may optionally be pairwise encoded by the stereo conversion component 910. For example, the stereo conversion component 910 may encode certain pairs of the KM output audio signals by performing MS or enhanced MS encoding.

[0109] The M output audio signals 922 (and any additional signals resulting from additional input audio signals 925) and the KM output audio signals 924 (or audio signals output from the stereo encoding component 910) are quantized by the quantization and multiplexing component 902 and included in the data stream 920. Additionally, parameters extracted by the various encoding components and modules may be quantized and included in the data stream.

[0110] The stereo encoding module 906 can operate in at least two configurations depending on the data transmission rate (bit rate) at which the encoder / decoder system operates, i.e., the bit rate at which the encoder 900 transmits data. The first configuration may correspond, for example, to a medium bit rate. The second configuration may correspond, for example, to a high bit rate. The encoder 900 includes an indication in the data stream 920 as to which configuration to use. For example, such an indication may be signaled via one or more bits in the data stream 920.

[0111] 10 illustrates the stereo encoding module 906 when operating according to a first configuration corresponding to a medium bit rate. The stereo encoding module 906 includes a first stereo conversion component 1040, various time / frequency conversion components 1042, 1046, an HFR encoding component 1048, a parametric stereo encoding component 1052, and a waveform encoding component 1056. The stereo encoding module 906 may further include a second stereo conversion component 1043. The stereo encoding module 906 takes as inputs two of the input audio signals 928. The input audio signals 928 are assumed to be represented in the time domain.

[0112] The first stereo conversion component 1040 converts the input audio signal 928 into a mid / side representation by forming sums and differences based on the above, and thus outputs a mid signal 1026 and a side signal 1024.

[0113] In some embodiments, the mid signal 1026 and the side signal 1024 are then converted to a mid / complement / a representation by a second stereo conversion component 1043. The second stereo conversion component 1043 extracts a weighting parameter a for inclusion in the data stream 920. The weighting parameter a may be time and frequency dependent, i.e., it may vary between different time frames and frequency bands of data.

[0114] The waveform encoding component 1056 subjects the mid signal 1026 and the side or complementary signal to waveform encoding, thereby generating a waveform-encoded mid signal 926 and a waveform-encoded side or complementary signal 924 .

[0115] The second stereo transform component 1043 and waveform coding component 1056 typically operate in the MDCT domain. Thus, the mid signal 1026 and the side signal 1024 may be transformed into the MDCT domain by the time-to-frequency transform component 1042 prior to the second stereo transform and waveform coding. If the signals 1026 and 1024 are not subjected to the second stereo transform 1043, different MDCT transform sizes may be used for the mid signal 1026 and the side signal 1024. If the signals 1026 and 1024 are subjected to the second stereo transform 1043, the same MDCT transform size should be used for the mid signal 1026 and the complementary signal 1024.

[0116] To achieve a moderate bit rate, the bandwidth of at least the side or complementary signal 924 is limited. More precisely, the side or complementary signal is waveform-encoded for frequencies up to a first frequency k1. Thus, the waveform-encoded side or complementary signal 924 includes spectral data corresponding to frequencies up to the first frequency k1. The mid signal 1026 is waveform-encoded for frequencies up to a frequency greater than the first frequency k1. Thus, the mid signal 926 includes spectral data corresponding to frequencies up to a frequency greater than the first frequency k1. In some cases, to save additional bits that need to be sent in the data stream 920, the bandwidth of the mid signal 926 is also limited. Thus, the waveform-encoded mid signal 926 includes spectral data up to a second frequency k2 greater than the first frequency k1.

[0117] If the mid signal 926 is bandwidth-limited, i.e., if the spectral content of the mid signal 926 is constrained to frequencies up to the second frequency k2, the mid signal 1026 is subjected to HFR encoding by the HFR encoding component 1048. Generally, the HFR encoding component 1048 analyzes the spectral content of the mid signal 1026 and extracts a set of parameters 1060. These parameters enable reconstruction of the spectral content of the signal for high frequencies (in this case, frequencies above the second frequency k2) based on the spectral content of the signal for low frequencies (in this case, frequencies above the second frequency k2). Such HFR encoding techniques are known in the art and include, for example, the spectral band replication (SBR) technique. The set of parameters 1060 is included in the data stream 920.

[0118] The HFR encoding component 1048 typically operates in the quadrature mirror filter (QMF) domain, so prior to performing HFR encoding, the mid signal 1026 may be converted to the QMF domain by the time-to-frequency transform component 1046.

[0119] The input audio signal 928 (or alternatively the mid signal 1046 and the side signal 1024) is subjected to parametric stereo encoding in a parametric stereo (PS) encoding component 1052. In general, the parametric stereo encoding component 1052 analyzes the input audio signal 928 and extracts parameters 1062 that allow reconstruction of the input audio signal 928 based on the mid signal 1026 for frequencies above a first frequency k1. The parametric stereo encoding component 1052 may apply any known technique for parametric stereo encoding.

[0120] The parametric stereo encoding component 1052 typically operates in the QMF domain, so the input audio signal 928 (or alternatively the mid signal 1046 and the side signal 1024) may be converted to the QMF domain by the time-to-frequency transform component 1046.

[0121] 11 illustrates the stereo encoding module 906 when functioning according to a second configuration supporting higher bit rates. The stereo encoding module 906 includes a first stereo conversion component 1140, various time-to-frequency conversion components 1142, 1146, HFR encoding components 1048a, 1048b, and a waveform encoding component 1156. Optionally, the stereo encoding module 906 may include a second stereo conversion component 1143. The stereo encoding module 906 takes as inputs two of the input audio signals 928. It is assumed that the input audio signals 928 are represented in the time domain.

[0122] The first stereo conversion component 1140 is similar to the first stereo conversion component 1040 and converts the input audio signal 928 into a mid signal 1126 and a side signal 1124 .

[0123] In some embodiments, the mid signal 1126 and the side signal 1124 are then converted to a mid / complement / a representation by a second stereo conversion component 1143. The second stereo conversion component 1143 extracts a weighting parameter a for inclusion in the data stream 920. The weighting parameter a may be time and frequency dependent; that is, it may vary between different time frames and frequency bands of data. A waveform coding component 1156 then subjects the mid signal 1126 and the side or complementary signal to waveform coding, thereby generating a waveform-coded mid signal 926 and a waveform-coded side or complementary signal 924.

[0124] The waveform encoding component 1156 is similar to the waveform encoding component 1056 of FIG. 10, with one important difference being the bandwidth of the output signals 926, 924. More precisely, the waveform encoding component 1156 performs waveform encoding of the mid signal 1126 and the side or complementary signal up to a second frequency k2 (which is typically greater than the first frequency k1 described for the medium rate case). As a result, the waveform-encoded mid signal 926 and the waveform-encoded side or complementary signal 924 contain spectral data corresponding to frequencies up to the second frequency k2. In some cases, the second frequency k2 may correspond to the maximum frequency represented by the system. In other cases, the second frequency k2 may be lower than the maximum frequency represented by the system.

[0125] If the second frequency k2 is lower than the maximum frequency represented by the system, the input audio signal 928 is subjected to HFR encoding by HFR components 1148a, 1148b. Each of the HFR encoding components 1148a, 1148b operates similarly to the HFR encoding component 1048 of FIG. 10. Thus, the HFR encoding components 1148a, 1148b generate a first set of parameters 1160a and a second set of parameters 1160b, respectively, which enable reconstruction of the spectral content of the respective input audio signal for high frequencies (in this case, frequencies above the second frequency k2) based on the spectral content of the input audio signal 928 for low frequencies (in this case, frequencies above the second frequency k2). The first and second sets of parameters 1160a, 1160b are included in the data stream 920.

[0126] Equivalents, extensions, replacements, etc. Further embodiments of the present disclosure will be apparent to those skilled in the art upon reviewing the above description. While the present text and drawings disclose embodiments and examples, the present disclosure is not limited to these specific examples. Numerous modifications and variations can be made without departing from the scope of the present disclosure, which is defined by the appended claims. Any reference signs appearing in the claims should not be construed as limiting the scope thereof.

[0127] Furthermore, variations to the disclosed embodiments can be understood and implemented by those skilled in the art in practicing the present disclosure, from an examination of the drawings, the disclosure, and the appended claims. In the claims, the word "comprises" does not exclude other elements or steps, and the singular does not exclude a plurality. The mere fact that certain features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.

[0128] The systems and methods disclosed above may be implemented as software, firmware, hardware, or a combination thereof. In hardware implementations, the division of tasks among functional units referred to in the above description does not necessarily correspond to a division into physical units. Conversely, a single physical component may have multiple functions, and a single task may be performed by several cooperating physical components. Some or all of the components may be implemented as software executed by a digital signal processor or microprocessor, or as hardware or application-specific integrated circuits. Such software may be distributed on computer-readable media, which may include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and that can be accessed by a computer. Additionally, those skilled in the art will recognize that communication media typically embodies computer-readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.

[0129] All drawings are schematic and generally show only parts necessary for clarity of the present disclosure, while other parts may be omitted or only suggested. Unless otherwise specified, like reference numerals refer to like parts in different drawings.

[0130] Several aspects will be described. [Aspect 1] 1. A method in a decoder for decoding a plurality of input audio signals for playback on a speaker configuration having N channels, the plurality of input audio signals representing encoded multi-channel audio content corresponding to at least N channels, the method comprising: receiving M input audio signals, <m≦n≦2mである、段階と;decoding, in a first decoding module, the M input audio signals into M mid signals suitable for playback on a speaker configuration having M channels; For each of more than M channels of the N channels, receiving an additional input audio signal corresponding to one of the M mid signals, the additional input audio signal being a side signal or a complementary signal that allows reconstruction of the side signal together with the mid signal and a weighting parameter a; decoding, in a stereo decode module, the additional input audio signal and its corresponding mid signal to generate a stereo signal including first and second audio signals suitable for playback on two of the N channels of the speaker configuration; thereby generating N audio signals suitable for playback on the N channels of said speaker configuration. method. [Aspect 2] The method of claim 1, wherein the stereo decode module is operable in at least two configurations depending on the bit rate at which the decoder receives data, and the method further includes receiving an instruction regarding which of the at least two configurations to use in decoding the additional input audio signal and its corresponding mid signal. Aspect 3 The step of receiving an additional input audio signal comprises: receiving a pair of audio signals corresponding to joint encoding of an additional input audio signal corresponding to a first of the M mid signals and an additional input audio signal corresponding to a second of the M mid signals; decoding the pair of audio signals to generate the additional input audio signals corresponding to the first and second of the M mid signals, respectively; The method of embodiment 1 or 2. Aspect 4 the additional input audio signal is a waveform-encoded signal including spectral data corresponding to frequencies up to a first frequency, and the corresponding mid signal is a waveform-encoded signal including spectral data corresponding to frequencies up to a frequency greater than the first frequency, and decoding the additional input audio signal and its corresponding mid signal in accordance with the first configuration of the stereo decoding module comprises: if the additional audio input signal is in the form of a complementary signal, calculating a side signal for frequencies up to the first frequency by multiplying the mid signal by a weighting parameter a and adding the result of the multiplication to the complementary signal; upmixing the mid signal and the side signal to generate a stereo signal including first and second audio signals, wherein for frequencies below the first frequency, the upmixing comprises performing an inverse sum-difference transform of the mid signal and the side signal, and for frequencies above the first frequency, the upmixing comprises performing a parametric upmix of the mid signal. The method of embodiment 2 or 3. Aspect 5 the waveform-encoded mid-signal includes spectral data corresponding to frequencies up to a second frequency, and the method further comprises: extending the mid signal to a frequency range above the second frequency by performing high frequency reconstruction prior to performing a parametric upmix. The method of embodiment 4. Aspect 6 the additional input audio signal and the corresponding mid signal are waveform-encoded signals including spectral data corresponding to frequencies up to a second frequency, and decoding the additional input audio signal and its corresponding mid signal in accordance with the second configuration of the stereo decoding module comprises: if the additional audio input signal is in the form of a complementary signal, calculating a side signal by multiplying the mid signal by the weighting parameter a and adding the result of the multiplication to the complementary signal; performing an inverse sum-difference transform of the mid signal and the side signal to generate a stereo signal including first and second audio signals. The method of embodiment 2 or 3. Aspect 7 further comprising extending the first and second audio signals of the stereo signal to a frequency range above the second frequency by performing high frequency reconstruction. The method of embodiment 6. Aspect 8 If the M mid signals are to be reproduced on a speaker configuration having M channels, the method further comprises: 8. The method of any one of aspects 1 to 7, further comprising: extending a frequency range of at least one of the M mid signals by performing high-frequency reconstruction based on high-frequency reconstruction parameters associated with the first and second audio signals of the stereo signal that may be generated from at least one of the M mid signals and its corresponding additional audio input signal. Aspect 9 9. The method of any one of aspects 1 to 8, wherein if the additional input audio signal is in the form of a side signal, the additional input audio signal and the corresponding mid signal are waveform-coded using a modified discrete cosine transform having different transform sizes. Aspect 10 10. A computer program product having a computer-readable medium with instructions for performing the method of any one of aspects 1 to 9. Aspect 11 1. A decoder for decoding a plurality of input audio signals for playback on a speaker configuration having N channels, the plurality of input audio signals representing encoded multi-channel audio content corresponding to at least N channels, the decoder comprising: a receiving component configured to receive M input audio signals, <m≦n≦2mである、受領コンポーネントと;a first decoding module configured to decode the M input audio signals into M mid signals suitable for playback on a speaker configuration having M channels; and a stereo encoding module for each of more than M channels of the N channels, the stereo encoding module comprising: receiving an additional input audio signal corresponding to one of the M mid signals, the additional input audio signal being a side signal or a complementary signal that allows reconstruction of the side signal together with the mid signal and a weighting parameter a; configured to decode the additional input audio signal and its corresponding mid signal to generate a stereo signal comprising first and second audio signals suitable for playback on two of the N channels of the speaker configuration; whereby the decoder is configured to generate N audio signals suitable for playback on the N channels of said speaker configuration. decoder. Aspect 12 1. A method in an encoder for encoding a plurality of input audio signals representing multi-channel audio content corresponding to K channels, comprising: receiving K input audio signals corresponding to channels of a speaker configuration having K channels; generating, from the K input audio signals, M mid signals and KM output audio signals suitable for playback on a speaker configuration having M channels, <m<k≦2mであり、2M-K of the mid signals correspond to 2M-K of the input audio signals; The remaining KM mid signals and the KM output audio signals are, for each value of K greater than M, a stereo encoding module for encoding two of the K input audio signals to generate a mid signal and an output audio signal, the output audio signal being a side signal or a complementary signal that allows reconstruction of the side signal together with the mid signal and a weighting parameter a; encoding the M mid signals into M additional output audio channels in a second encoding module; and including the KM output audio signals and the M additional output audio channels in a data stream for transmission to a decoder. method. Aspect 13 13. The method of claim 12, wherein the stereo encoding module is operable in at least two configurations depending on a desired bit rate of the encoder, and the method further comprises including in the data stream an indication as to which of the at least two configurations was used by the stereo encoding module in encoding two of the K input audio signals. Aspect 14 14. The method of claim 12 or 13, further comprising: performing stereo encoding of the KM output audio signals for each pair prior to including them in the data stream. Aspect 15 When the stereo encoding module operates according to a first configuration, encoding two of the K input audio signals to generate a mid signal and an output audio signal includes: converting the two input audio signals into a first signal that is a mid signal and a second signal that is a side signal; waveform encoding the first and second signals into first and second waveform-encoded signals, respectively, wherein the second signal is waveform-encoded to a first frequency and the first signal is waveform-encoded to a second frequency greater than the first frequency; subjecting the two input audio signals to parametric stereo encoding to extract parametric stereo parameters that allow reconstruction of spectral data of the two of the K input audio signals for frequencies above the first frequency; and including said first and second waveform-encoded signals and said parametric stereo parameters in said data stream. 15. The method of any one of embodiments 12 to 14. Aspect 16 for frequencies below the first frequency, converting the waveform-coded second signal, which is a side signal, into a complementary signal by multiplying the waveform-coded first signal, which is a mid signal, by a weighting factor a and subtracting the result of the multiplication from the second waveform-coded signal; and including the weighting parameter a in the data stream. The method of embodiment 15. Aspect 17 subjecting the first signal, being a mid-signal, to high-frequency reconstruction encoding to generate high-frequency reconstruction parameters that enable high-frequency reconstruction of the first signal above the second frequency; and including the high-frequency reconstruction parameters in the data stream. 17. The method of embodiment 15 or 16. Aspect 18 When the stereo encoding module operates according to a second configuration, encoding two of the K input audio signals to generate a mid signal and an output audio signal includes: converting the two input audio signals into a first signal that is a mid signal and a second signal that is a side signal; waveform encoding the first and second signals into first and second waveform-encoded signals, respectively, wherein the first and second signals are waveform-encoded to a second frequency; including the first and second waveform-encoded signals. 15. The method of any one of embodiments 12 to 14. Aspect 19 converting the waveform-coded second signal, a side signal, into a complementary signal by multiplying the waveform-coded first signal, a mid signal, by a weighting factor a and subtracting the result of the multiplication from the second waveform-coded signal; and including the weighting parameter a in the data stream. The method of embodiment 18. Aspect 20 subjecting each of the two of the K input audio signals to high-frequency reconstruction encoding to generate high-frequency reconstruction parameters that enable high-frequency reconstruction of the two of the N input audio signals above the second frequency; and including the high-frequency reconstruction parameters in the data stream. 20. The method of embodiment 18 or 19. Aspect 21 21. A computer program product having a computer-readable medium having instructions for performing the method of any one of aspects 12 to 20. Aspect 22 1. An encoder for encoding a plurality of input audio signals representing multi-channel audio content corresponding to K channels, comprising: a receiving component configured to receive K input audio signals corresponding to channels of a speaker configuration having K channels; a first encoding module configured to generate, from the K input audio signals, M mid signals and KM output audio signals suitable for playback on a speaker configuration having M channels, the first encoding module comprising: <m<k≦2mであり、2M-K of the mid signals correspond to 2M-K of the input audio signals; The first encoding module includes KM stereo encoding modules configured to generate the remaining KM mid signals and KM output audio signals, each stereo encoding module comprising: a first encoding module configured to encode two of the K input audio signals to generate a mid signal and an output audio signal, the output audio signal being a side signal or a complementary signal that allows reconstruction of the side signal together with the mid signal and a weighting parameter a; a second encoding module configured to encode the M mid signals into M additional output audio channels; a multiplexing component configured to include the KM output audio signals and the M additional output audio channels in a data stream for transmission to a decoder. Encoder.

Claims

1. 1. A method for decoding a plurality of audio signals, the method comprising: determining a stereo signal based on a mid signal and a side signal; the plurality of audio signals include the mid signal and the side signal; the stereo signal includes a first audio signal and a second audio signal suitable for playback on two channels of a speaker configuration; the stereo signal is determined based on a first up-mix comprising performing a weighted inverse sum-difference transform of the mid signal and the side signal for a first frequency below the first frequency, and a second up-mix comprising an up-mix of the mid signal for a second frequency above the first frequency; method.

2. The first audio signal includes spectral data corresponding to a third frequency through a second frequency, and the method further comprises: extending the first audio signal to a frequency range above the second frequency by performing high frequency reconstruction before performing a parametric upmix. The method of claim 1.

3. A non-transitory computer-readable storage medium containing instructions that, when executed by a processor, perform the method of claim 1.

4. 1. An apparatus for decoding a plurality of audio signals, the apparatus comprising: a processor for determining a stereo signal based on the mid signal and the side signal, the plurality of audio signals include the mid signal and the side signal; the stereo signal includes a first audio signal and a second audio signal suitable for playback on two channels of a speaker configuration; the stereo signal is determined based on a first up-mix comprising performing a weighted inverse sum-difference transform of the mid signal and the side signal for a first frequency below the first frequency, and a second up-mix comprising an up-mix of the mid signal for a second frequency above the first frequency; Device.