Audio decoder for interleaving a signal

The hybrid coding method improves multi-channel audio coding efficiency by waveform-coding low frequencies and parametrically restoring high frequencies, addressing bandwidth inefficiencies and quality limitations in conventional methods.

JP7778768B2Active Publication Date: 2025-12-02DOLBY INTERNATIONAL AB
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023220177
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2013-04-05
Filing Date
2023-12-27
Publication Date
2025-12-02
Estimated Expiration
2034-04-04

AI Technical Summary

Technical Problem

Existing multi-channel audio coding technologies face inefficiencies in bandwidth utilization and quality at bit-rates between low and high, particularly with conventional parametric coding saturating at around 72 kbps and discrete coding not fully leveraging human auditory sensitivity to low frequencies.

Method used

A hybrid coding approach combining parametric and discrete multi-channel coding, where low frequencies are waveform-coded with higher frequencies restored parametrically, allowing for improved quality and reduced bit-rate by using separate crossover frequencies for different coding methods.

Benefits of technology

Enhances audio quality by leveraging human auditory sensitivity to low frequencies and optimizing bit allocation, reducing overall bit-rate requirements while maintaining or improving perceived audio quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007778768000001
    Figure 0007778768000001
  • Figure 0007778768000002
    Figure 0007778768000002
  • Figure 0007778768000003
    Figure 0007778768000003
Patent Text Reader

Abstract

To provide a method and device for decoding an audio bitstream encoded by an audio processing system.SOLUTION: A decoder 100 of a multi-channel audio processing system receives N waveform encoded down-mix signals and M waveform encoded signals representing a multi-channel audio signal to be decoded, wherein a first conceptual element 200 and the M waveform-encoded signals that satisfy 1<N<M are down-mixed and then combined with the N waveform encoded down-mix signals, a second conceptual element 300 and a high-frequency restored signal on which high-frequency restoration (HFR) is executed for the combined down-mix signals are up-mixed, and the M waveform encoded signals include a third conceptual element 400 combined with up-mix signals so as to restore M encoded channels.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The disclosure herein relates generally to multi-channel audio coding. In particular, the disclosure relates to encoders and decoders for hybrid coding, including parametric coding and discrete multi-channel coding.

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of U.S. Patent Application No. 14 / 772,001, filed September 1, 2015, which is a Section 371 national stage application of PCT Application No. PCT / EP2014 / 056852, filed April 4, 2014, which in turn claims priority to U.S. Provisional Patent Application No. 61 / 808,680, filed April 5, 2013, and each of these applications is hereby incorporated by reference in its entirety. [Background technology]

[0003] In conventional multi-channel audio coding, possible coding schemes include discrete multi-channel coding, such as MPEG Surround®, or parametric coding. The scheme used depends on the bandwidth of the audio system. Parametric coding methods are known to be scalable and efficient in terms of listening quality, making them particularly attractive for low bit-rate applications. For high bit-rate applications, discrete multi-channel coding is often used. Especially for applications with bit-rates between low and high bit-rates, existing distribution or processing formats and accompanying coding techniques can be improved in terms of their bandwidth efficiency. Summary of the Invention [Problem to be solved by the invention]

[0004] U.S. Patent No. 7,292,901 (US7,292,901) (Kroon et al.) relates to a hybrid coding method in which a hybrid audio signal is formed from at least one downmixed spectral component and at least one pure (unmixed) spectral component. The method disclosed therein can increase the capacity of an application having a certain bit rate, but further improvements may be needed to further increase the efficiency of audio processing systems.

[0005] Example embodiments will now be described with reference to the accompanying drawings, in which: [Brief explanation of the drawings]

[0006] [Figure 1] FIG. 1 is a generalized block diagram of a decoding system according to an example embodiment. [Figure 2] FIG. 2 illustrates a first part of the decoding system in FIG. 1. [Figure 3] FIG. 2 illustrates a second part of the decoding system in FIG. 1. [Figure 4] FIG. 2 illustrates a third part of the decoding system in FIG. 1. [Figure 5] FIG. 1 is a generalized block diagram of an encoding system according to an example embodiment. [Figure 6] FIG. 1 is a generalized block diagram of a decoding system according to an example embodiment. [Figure 7] FIG. 7 illustrates a third part of the decoding system in FIG. 6. [Figure 8] FIG. 1 is a generalized block diagram of an encoding system according to an example embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0007] All drawings are schematic and generally show only the elements necessary to explain the present disclosure, while other elements may be omitted or merely suggested. Unless otherwise indicated, like reference numerals refer to like elements in different drawings.

[0008] "Overview of the Decoder" As used herein, an audio signal can be a pure audio signal, the audio portion of an audiovisual or multimedia signal, or any of these combined with metadata.

[0009] As used herein, downmixing of a plurality of signals means combining the plurality of signals, for example, by forming a primary combination so that a smaller number of signals are obtained. The reverse operation to downmixing is called upmixing, that is, operating on a smaller number of signals so that a larger number of signals are obtained.

[0010] According to a first aspect, example embodiments propose a method, an apparatus, and a computer program product for restoring a multi-channel audio signal based on an input signal. The proposed method, apparatus, and computer program product may generally have the same features and advantages.

[0011] According to an example embodiment, a decoder suitable for a multi-channel audio processing system for restoring M (M>2) encoded channels is provided. The decoder includes a first receiving stage configured to receive N (1<N<M) waveform-encoded downmix signals including spectral coefficients corresponding to frequencies between a first crossover frequency and a second crossover frequency.

[0012] The decoder further includes a second receiving stage configured to receive M waveform-encoded signals including spectral coefficients corresponding to frequencies up to the first crossover frequency, each of the M waveform-encoded signals corresponding to a respective one of the M encoded channels.

[0013] The decoder further includes a downmix stage downstream of the second receiving stage configured to downmix the M waveform-encoded signals into N downmix signals including spectral coefficients corresponding to frequencies up to the first crossover frequency.

[0014] The decoder further includes a first combining stage downstream of the first receiving stage and the downmix stage configured to combine each of the N waveform-coded downmix signals received by the first receiving stage with a corresponding one of the N downmix signals from the downmix stage into N combined downmix signals.

[0015] The decoder further includes a high-frequency restoration stage downstream of the first combining stage configured to extend each of the N combined downmix signals from the first combining stage to a frequency range above the second crossover frequency by performing high-frequency restoration.

[0016] The decoder further includes an upmix stage downstream of the high-frequency restoration stage configured to perform a parametric upmix of the frequency-extended N combined downmix signals from the high-frequency restoration stage into M upmix signals comprising spectral coefficients corresponding to frequencies above the first crossover frequency, each of the M upmix signals corresponding to one of the M encoded channels.

[0017] The decoder further includes a second combining stage downstream of the upmix stage and the second receiving stage configured to combine the M upmix signals from the upmix stage with the M waveform-encoded signals received by the second receiving stage.

[0018] The M waveform-coded signals are pure waveform-coded signals without any parametric signals mixed in, i.e., they are non-downmixed discrete representations of the processed multi-channel audio signal. An advantage of having lower frequencies represented in these waveform-coded signals may be that the human ear is more sensitive to parts of an audio signal that have low frequencies. By encoding this part with better quality, the overall impression of the decoded audio may be enhanced.

[0019] An advantage of having at least two downmix signals is that this embodiment provides increased dimensionality of the downmix signal compared to a system with only one downmix channel, and therefore may provide better decoded audio quality, which may exceed the gain in bit rate provided by a one downmix signal system.

[0020] The advantage of using hybrid coding including parametric downmixing and discrete multi-channel coding is that it can improve the quality of the decoded audio signal for a certain bitrate compared to a conventional parametric coding approach, i.e., MPEG Surround with HE-AAC. At bitrates of approximately 72 kilobits per second (kbps), the conventional parametric coding model may saturate, i.e., the quality of the decoded audio signal is limited not by a lack of bits for coding but by the shortcomings of the parametric model. Therefore, for bitrates from approximately 72 kbps, it may be more beneficial to use bits to discretely waveform code lower frequencies. At the same time, the hybrid approach using parametric downmixing and discrete multi-channel coding can improve the quality of the decoded audio signal for a certain bitrate, for example, 128 kbps or less, compared to using an approach in which all bits are used to waveform code lower frequencies and using spectral band replication (SBR) for the remaining frequencies.

[0021] An advantage of having N waveform-coded downmix signals that include only spectral data corresponding to frequencies between the first and second crossover frequencies is that the required bit rate for the audio signal processing system can be reduced. Alternatively, the bits saved by having a bandpass filtered downmix signal can be used to waveform-code lower frequencies, for example, the sample frequency for those frequencies can be made higher or the first crossover frequency can be increased.

[0022] As mentioned above, since the human ear is more sensitive to portions of an audio signal having low frequencies, high frequencies as portions of the audio signal having frequencies above the second crossover frequency can be reproduced by high frequency restoration without reducing the perceived audio quality of the decoded audio signal.

[0023] A further advantage of this embodiment may be that the parametric upmix performed in the upmix stage processes only spectral coefficients corresponding to frequencies above the first crossover frequency, thereby reducing the complexity of the upmix.

[0024] According to another embodiment, the combining performed in the first combining stage, in which each of the N waveform-coded downmix signals comprising spectral coefficients corresponding to frequencies between a first crossover frequency and a second crossover frequency is combined with a corresponding one of the N downmix signals comprising spectral coefficients corresponding to frequencies up to the first crossover frequency to obtain N combined downmix signals, is performed in the frequency domain.

[0025] An advantage of this embodiment may be that the M waveform-coded signals and the N waveform-coded downmix signals may be coded by a waveform coder using overlapping windowed transforms with independent windowing on the M waveform-coded signals and the N waveform-coded downmix signals, respectively, and may still be decodable by a decoder.

[0026] According to another embodiment, extending each of the N combined downmix signals in the high frequency restoration stage to a frequency range above the second crossover frequency is performed in the frequency domain.

[0027] According to a further embodiment, the combining performed in the second combining stage, i.e. the combining of the M upmix signals comprising spectral coefficients corresponding to frequencies above the first crossover frequency with the M waveform-coded signals comprising spectral coefficients corresponding to frequencies up to the first crossover frequency, is performed in the frequency domain. As mentioned above, an advantage of combining signals in the QMF domain is that an independent windowing of the overlapping windowed transform used to encode the signal in the MDCT domain can be used.

[0028] According to another embodiment, the parametric upmixing of the frequency-extended N combined downmix signals into M upmix signals, performed in the upmix stage, is performed in the frequency domain.

[0029] According to yet another embodiment, downmixing the M waveform-coded signals into N downmix signals comprising spectral coefficients corresponding to frequencies up to the first crossover frequency is performed in the frequency domain.

[0030] According to one embodiment, the frequency domain is the Quadrature Mirror Filter (QMF) domain.

[0031] According to another embodiment, the downmixing performed in a downmixing stage, in which the M waveform-coded signals are downmixed into N downmix signals comprising spectral coefficients corresponding to frequencies up to a first crossover frequency, is performed in the time domain.

[0032] According to yet another embodiment, the first crossover frequency is determined by the bit rate of the multi-channel audio processing system, which may result in that parts of the audio signal having frequencies below the first crossover frequency are simply waveform encoded, so that the available bandwidth is utilized to improve the quality of the decoded audio signal.

[0033] According to another embodiment, extending each of the N combined downmix signals to a frequency range above the second crossover frequency by performing high frequency restoration in a high frequency restoration stage is performed using high frequency restoration parameters. The high frequency restoration parameters may be received by the decoder, for example in a receiving stage, and then transmitted to the high frequency restoration stage. The high frequency restoration may include, for example, performing Spectral Band Replication (SBR).

[0034] According to another embodiment, the parametric upmix in the upmixing stage is performed with the use of upmix parameters, which are received by the encoder, for example in a receiving stage, and transmitted to the upmixing stage. Decorrelated versions of the N frequency-extended combined downmix signals are generated, and a matrix operation is performed on the N frequency-extended combined downmix signals and on the decorrelated versions of the N frequency-extended combined downmix signals, the parameters of which are given by the upmix parameters.

[0035] According to another embodiment, the received N waveform-coded downmix signals at the first receiving stage and the received M waveform-coded signals at the second receiving stage are coded using overlap windowed transforms with independent windowing for the N waveform-coded downmix signals and the M waveform-coded signals, respectively.

[0036] An advantage of this may be that it allows for improved coding quality and therefore an increased quality of the decoded multi-channel audio signal. For example, if at a certain point in time a transient signal is detected in a higher frequency band, the waveform encoder may encode this special time frame with a shorter window sequence, while the default window sequence for the lower frequency bands may be retained.

[0037] According to an embodiment, the decoder may include a third receiving stage configured to receive a further waveform-encoded signal including spectral coefficients corresponding to a subset of frequencies above the first crossover frequency. The decoder may further include an interleaving stage downstream of the upmix stage. The interleaving stage may be configured to interleave the further waveform-encoded signal with one of the M upmix signals. The third receiving stage may be further configured to receive a plurality of further waveform-encoded signals, and the interleaving stage may be further configured to interleave the plurality of further waveform-encoded signals with the plurality of M upmix signals.

[0038] This is advantageous in that certain parts of the frequency range above the first crossover frequency that are difficult to parametrically reconstruct from the downmix signal can be provided in waveform-coded form as a result of interleaving with the parametrically reconstructed upmix signal.

[0039] In one exemplary embodiment, the interleaving is performed by adding the further waveform-coded signal with one of the M upmix signals. According to another exemplary embodiment, the step of interleaving the further waveform-coded signal with one of the M upmix signals comprises replacing one of the M upmix signals by the further waveform-coded signal at a subset of frequencies above the first crossover frequency corresponding to spectral coefficients of the further waveform-coded signal.

[0040] According to an exemplary embodiment, the decoder may further be configured to receive a control signal, for example by a third receiving stage. The control signal may indicate how to interleave the further waveform-coded signal with one of the M upmix signals, and the step of interleaving the further waveform-coded signal with one of the M upmix signals is based on the control signal. Specifically, the control signal may indicate a frequency range and a time range, such as one or more time / frequency tiles in the QMF domain, in which the further waveform-coded signal should be interleaved with one of the M upmix signals. Thus, interleaving may occur in time and frequency within one channel.

[0041] The advantage of this is that time and frequency ranges can be selected that do not suffer from the aliasing or start-up / fade-out problems of the overlapping windowed transform used to encode the waveform-coded signal.

[0042] According to some embodiments, a method for decoding an encoded audio bitstream in an audio processing system is disclosed. The method includes extracting from the encoded audio bitstream a first waveform-coded signal including spectral coefficients corresponding to frequencies up to a first crossover frequency, and performing parametric decoding at a second crossover frequency to generate a reconstructed signal. The second crossover frequency is above the first crossover frequency, and the parametric decoding generates the reconstructed signal using reconstruction parameters obtained from the encoded audio bitstream. The method further includes extracting from the encoded audio bitstream a second waveform-coded signal including spectral coefficients corresponding to a subset of frequencies above the first crossover frequency, and interleaving the second waveform-coded signal with the reconstructed signal to generate an interleaved signal. The interleaved signal is then combined with the first waveform-coded signal.

[0043] Many variations exist as well. For example, the first crossover frequency may depend on the bit rate of the audio processing system, and the interleaving step may include (i) adding the second waveform-coded signal with the reconstructed signal, (ii) combining the second waveform-coded signal with the reconstructed signal, or (iii) replacing the reconstructed signal with the second waveform-coded signal. The combining step of the interleaved signal with the first waveform-coded signal may be performed in the frequency domain, or the performing of parametric decoding at the second crossover frequency to generate the reconstructed signal may be performed in the frequency domain. The parametric decoding may include either (i) parametric upmixing using upmix parameters, or (ii) high-frequency restoration using high-frequency restoration parameters, such as spectral band replication (SBR). The method may further include receiving a control signal used during the interleaving step to generate the interleaved signal. The control signal may indicate how to interleave the second waveform-coded signal with the restored signal by specifying either a frequency range or a time range for the interleaving step. A first value of the control signal may indicate that the interleaving step is to be performed for each frequency range. The interleaving step may similarly be performed before the combining step. The interleaving step and the combining step may similarly be combined into a single stage or operation. The first waveform-coded signal and the second waveform-coded signal may comprise signals representing the waveform of the audio signal in the frequency or time domain.

[0044] "Encoder Overview" According to a second aspect, illustrative embodiments propose a method, an apparatus, and a computer program product for encoding a multi-channel audio signal based on an input signal.

[0045] The proposed methods, apparatus, and computer program products may generally have the same features and advantages.

[0046] The advantages related to the features and configurations presented in the overview of the above decoder can generally be effective for the corresponding features and configurations for the encoder.

[0047] According to an embodiment of the example, an encoder suitable for a multi-channel audio processing system for encoding M (M>2) channels is provided.

[0048] The encoder includes a receiving stage configured to receive M signals corresponding to the M channels to be encoded.

[0049] The encoder further includes a first waveform encoding stage configured to receive the M signals from the receiving stage and generate M waveform-encoded signals including spectral coefficients corresponding to frequencies up to a first crossover frequency by individually waveform-encoding the M signals with respect to a frequency range corresponding to frequencies up to the first crossover frequency.

[0050] The encoder further includes a downmixing stage configured to receive the M signals from the receiving stage and downmix the M signals into N (1<N<M) downmix signals.

[0051] The encoder further includes a high-frequency restoration encoding stage configured to receive the N downmix signals from the downmixing stage and perform high-frequency restoration encoding on the N downmix signals, the high-frequency restoration encoding stage being configured to extract high-frequency restoration parameters that enable high-frequency restoration of the N downmix signals above a second crossover frequency.

[0052] The encoder further includes a parametric encoding stage configured to receive the M signals from the receiving stage and the N downmix signals from the downmixing stage and to perform parametric encoding on the M signals for a frequency range corresponding to frequencies above a first crossover frequency, the parametric encoding stage configured to extract upmix parameters that enable upmixing of the N downmix signals into M restored signals corresponding to the M channels for the frequency range above the first crossover frequency.

[0053] The encoder further includes a second waveform encoding stage configured to receive the N downmix signals from the downmixing stage and generate N waveform-coded downmix signals by waveform-coding the N downmix signals for a frequency range corresponding to frequencies between a first crossover frequency and a second crossover frequency, wherein the N waveform-coded downmix signals include spectral coefficients corresponding to frequencies between the first crossover frequency and the second crossover frequency.

[0054] According to one embodiment, performing high frequency reconstruction coding on the N downmix signals in the high frequency reconstruction coding stage is performed in the frequency domain, preferably in the quadrature mirror filter (QMF) domain.

[0055] According to a further embodiment, performing parametric coding on the M signals in the parametric coding stage is performed in the frequency domain, preferably in the quadrature mirror filter (QMF) domain.

[0056] According to yet another embodiment, generating M waveform-coded signals by individually waveform-coding the M signals in the first waveform-coding stage comprises applying an overlapping windowing transform to the M signals, wherein different overlapping window sequences are used for at least two of the M signals.

[0057] According to an embodiment, the encoder further comprises a third waveform encoding stage configured to generate a further waveform-encoded signal by waveform encoding one of the M signals for a frequency range corresponding to a subset of the frequency range above the first crossover frequency.

[0058] According to an embodiment, the encoder may include a control signal generation stage configured to generate a control signal indicating how the further waveform-coded signal should be interleaved with a parametric reconstruction of one of the M signals at the decoder. For example, the control signal may indicate a frequency range and a time range in which the further waveform-coded signal should be interleaved with one of the M upmix signals.

[0059] Illustrative Examples FIG. 1 is a generalized block diagram of a decoder 100 in a multi-channel audio processing system for restoring M encoded channels. The decoder 100 includes three conceptual elements 200, 300, 400 that will be described in more detail in connection with FIGS. 2 through 4. In a first conceptual element 200, the decoder receives N waveform-encoded downmix signals and M waveform-encoded signals representing the multi-channel audio signal to be decoded, where 1 < N < M. In the illustrated example, N is set to 2. In a second conceptual element 300, the M waveform-encoded signals are downmixed and combined with the N waveform-encoded downmix signals. High-frequency restoration (HFR) is then performed for the combined downmix signals. In a third conceptual element 400, the high-frequency restored signal is upmixed and the M waveform-encoded signals are combined with the upmix signal to restore the M encoded channels.

[0060] In the representative embodiments described in connection with FIGS. 2 through 4, the restoration of encoded 5. surround audio is described. It may be noted that the low frequency effect signal is not mentioned in the described embodiments or drawings. This does not mean that any low frequency effects are ignored. The low frequency effect (Lfe) is added to the five restored channels in any suitable manner well known to those skilled in the art. It may also be noted that the described decoder is equally well suited for other types of encoded surround audio, such as 7.1 or 9.1 surround audio.

[0061] Figure 2 illustrates a first conceptual element 200 of the decoder 100 in Figure 1. The decoder includes two receiving stages 212, 241. In the first receiving stage 212, the bitstream 202 is decoded and dequantized into two waveform-coded downmix signals 208a-b. Each of the two waveform-coded downmix signals 208a-b has a first crossover frequency k y and the second crossover frequency k x It contains spectral coefficients corresponding to frequencies between

[0062] In the second receiver stage 214, the bitstream 202 is decoded and dequantized into five waveform-encoded signals 210a-e, each of which has a first crossover frequency k y It contains spectral coefficients corresponding to frequencies up to

[0063] As an example, signals 210a-e include two channel pair components and one single-channel component for the center. The channel pair components can be, for example, a combination of a left front signal and a left surround signal, and a combination of a right front signal and a right surround signal. Further examples are a combination of a left front signal and a right front signal, and a combination of a left surround signal and a right surround signal. These channel pair components can be coded, for example, in a sum-and-difference format. All five signals 210a-e can be coded using overlapping windowed transforms with independent windowing and still be decodable by a decoder. This can enable improved coding quality and, therefore, improved quality of the decoded signals.

[0064] As an example, the first crossover frequency k y is 1.1 kHz. As an example, the second crossover frequency k x The first crossover frequency k is in the range of 5.6 to 8 kHz. ymay vary, even on an individual signal basis, i.e., the encoder can detect that a signal component in a particular output signal may not be faithfully reproduced by the stereo downmix signals 208a-b, and then adjust the bandwidth, i.e., the first crossover frequency k of the associated waveform-coded signal, i.e., 210a-e, during that particular time instance in order to perform appropriate waveform coding of the signal component. y It should be noted that it is possible to increase

[0065] As will be explained later in this description, the remaining stages of the decoder 100 generally operate in the Quadrature Mirror Filter (QMF) domain. For this reason, each of the signals 208a-b, 210a-e received by the first and second receiver stages 212, 214 in modified discrete cosine transform (MDCT) form is transformed into the time domain by applying an inverse MDCT 216. Each signal is then transformed back into the frequency domain by applying a QMF transform 218.

[0066] In FIG. 3, five waveform-encoded signals 210 are downmixed in a downmix stage 308 at a first crossover frequency k y The low-pass multi-channel signals 210a-e are downmixed into two downmix signals 310, 312 containing spectral coefficients corresponding to frequencies up to and including 100 kHz. These downmix signals 310, 312 may be formed by performing a downmix on the low-pass multi-channel signals 210a-e using the same downmixing scheme as was used in the encoder to create the two downmix signals 208a-b shown in FIG.

[0067] The two new downmix signals 310, 312 are then combined with the corresponding downmix signals 208a-b in a first combining stage 320, 322 to form combined downmix signals 302a-b. Each of the combined downmix signals 302a-b therefore has a first crossover frequency k , where the downmix signals 310, 312 originate. y and the spectral coefficients corresponding to frequencies up to the first crossover frequency k originating from the two waveform-coded downmix signals 208a-b received in the first receiving stage 212 (shown in FIG. 2). y and the second crossover frequency k x and spectral coefficients corresponding to frequencies between

[0068] The decoder further includes a high frequency reconstruction (HFR) stage 314. The HFR stage performs high frequency reconstruction to convert each of the two combined downmix signals 302a-b from the combining stage to a second crossover frequency k x The HFR stage 314 is configured to extend the frequency range above the HFR stage 314. According to some embodiments, the high frequency restoration performed includes performing spectral band replication (SBR). The high frequency restoration may be performed by using high frequency restoration parameters that may be received by the HFR stage 314 in any suitable manner.

[0069] The output from the high frequency restoration stage 314 are two signals 304a-b comprising the downmix signals 208a-b with applied HFR extension portions 316, 318. As explained above, the HFR stage 314 will perform high frequency restoration based on the frequencies present in the input signals 210a-e from the second receiving stage 214 (shown in FIG. 2) combined with the two downmix signals 208a-b. Somewhat simplified, the HFR ranges 316, 318 comprise the portions of the spectral coefficients from the downmix signals 310, 312 that have been copied down to the HFR ranges 316, 318. Thus, portions of the five waveform-encoded signals 210a-e will appear in the HFR ranges 316, 318 of the output 304 from the HFR stage 314.

[0070] It should be noted that the downmixing in the downmixing stage 308 prior to the high-frequency reconstruction stage 314 and the combining in the first combining stages 320, 322 can be performed in the time domain, i.e., after each signal is transformed into the time domain by applying the inverse modified discrete cosine transform (MDCT) 216 (shown in FIG. 2 ). However, if the waveform-coded signals 210a-e and the waveform-coded downmix signals 208a-b are likely to be coded by a waveform coder using overlapping windowed transforms with independent windowing, the signals 210a-e and 208a-b may not be seamlessly combined in the time domain. Therefore, a better-controlled scenario is achieved if the combining in at least the first combining stages 320, 322 is performed in the QMF domain.

[0071] 4 illustrates the third and final conceptual element 400 of the decoder 100. The output 304 from the HFR stage 314 forms the input to an upmix stage 402. The upmix stage 402 performs a parametric upmix on the frequency extended signals 304a-b, thereby producing five signal outputs 404a-e. Each of the five upmix signals 404a-e has a first crossover frequency k y The frequency-extended combined downmix signal 304a-b corresponds to one of the five encoded channels in the encoded 5.1 surround sound for higher frequencies. According to a typical parametric upmix procedure, the upmix stage 402 first receives parametric mixing parameters. The upmix stage 402 further generates a decorrelated version of the two frequency-extended combined downmix signals 304a-b. The upmix stage 402 further performs a matrix operation on the two frequency-extended combined downmix signals 304a-b and the decorrelated version of the two frequency-extended combined downmix signals 304a-b, where the parameters of the matrix operation are given by the upmix parameters. Alternatively, any other parametric upmix procedure known in the art may be applied. Applicable parametric upmixing procedures are described, for example, in "MPEG Surround - The ISO / MPEG Standard for Efficient and Compatible Multichannel Audio Coding" (Herre et al., Journal of the Audio Engineering Society, Vol. 56, No. 11, November 2008).

[0072] Thus, the outputs 404a-e from the upmix stage 402 are at the first crossover frequency k y Does not include frequencies below the first crossover frequency k yThe spectral coefficients corresponding to the remaining frequencies up to are present in the five waveform-encoded signals 210 a - e which have been delayed by delay stage 412 to match the timing of the upmix signal 404 .

[0073] The decoder 100 further includes second combining stages 416, 418. The second combining stages 416, 418 are configured to combine the five upmix signals 404a-e with the five waveform-encoded signals 210a-e received by the second receiving stage 214 (shown in FIG. 2).

[0074] It may be noted that any current Lfe signal can be added as a separate signal to the resulting combined signal 422. Each of the signals 422 is then transformed into the time domain by applying an inverse QMF transform 414. Thus, the output from the inverse QMF transform 414 is a fully decoded 5.1 channel audio signal.

[0075] Figure 6 illustrates a decoding system 100' that is an improved version of the decoding system 100. The decoding system 100' has conceptual elements 200', 300', and 400' that correspond to the conceptual elements 200, 300, and 400 of Figure 1. The difference between the decoding system 100' of Figure 6 and the decoding system of Figure 1 is the presence of a third receiving stage 616 in the conceptual element 200' and an interleaving stage 714 in the third conceptual element 400'.

[0076] The third receiving stage 616 is configured to receive a further waveform-encoded signal. The further waveform-encoded signal includes spectral coefficients corresponding to a subset of frequencies above the first crossover frequency. The further waveform-encoded signal may be transformed into the time domain by applying an inverse MDCT 216. It may then be transformed back into the frequency domain by applying a QMF transform 218.

[0077] It should be understood that the further waveform-encoded signal may be received as a separate signal. However, the further waveform-encoded signal may likewise form part of one or more of the five waveform-encoded signals 210a-e. In other words, the further waveform-encoded signal may be coded together with one or more of the five waveform-encoded signals 210a-e, for example using the same MDCT transform. If so, the third receiver stage 616 corresponds to the second receiver stage, i.e., the further waveform-encoded signal is received together with the five waveform-encoded signals 210a-e by the second receiver stage 214.

[0078] 7 illustrates in more detail the third conceptual element 300' of the decoder 100' of FIG. 6. In addition to the high-frequency extended downmix signals 304a-b and the five waveform-encoded signals 210a-e, a further waveform-encoded signal 710 is input to the third conceptual element 400'. In the illustrated example, the further waveform-encoded signal 710 corresponds to the third channel of the five channels. The further waveform-encoded signal 710 corresponds to the first crossover frequency k y , and further includes spectral coefficients corresponding to a frequency interval starting from . However, the form of the subset of the frequency range above the first crossover frequency covered by the further waveform-coded signal 710 may of course vary in different embodiments. It should also be noted that multiple waveform-coded signals 710a-e may be received, and different waveform-coded signals may correspond to different output channels. The subset of the frequency range covered by the multiple further waveform-coded signals 710a-e may vary between different ones of the multiple further waveform-coded signals 710a-e.

[0079] The further waveform-encoded signal 710 may be delayed by a delay stage 712 to match the timing of the upmix signal 404 output from the upmix stage 402. The upmix signal 404 and the further waveform-encoded signal 710 are then input to an interleaving stage 714. The interleaving stage 714 interleaves, i.e., combines, the upmix signal 404 with the further waveform-encoded signal 710 to generate an interleaved signal 704. In this example, the interleaving stage 714 thus interleaves the third upmix signal 404c with the further waveform-encoded signal 710. Interleaving may be performed by adding the two signals together. However, in general, interleaving is performed by swapping the upmix signal 404 with the further waveform-encoded signal 710 in the frequency and time ranges where the signals overlap.

[0080] The interleaved signal 704 is then input to the second combining stage 416, 418, where it is combined with the waveform-encoded signals 201a-e in the same manner as described with reference to Figure 4 to produce the output signal 722. It should be noted that the order of the interleaving stage 714 and the second combining stage 416, 418 may be reversed so that the combining occurs before the interleaving.

[0081] Furthermore, in situations where the further waveform-coded signal 710 forms part of one or more of the five waveform-coded signals 210a-e, the second combining stage 416, 418 and the interleaving stage 714 may be combined into a single stage. Specifically, such a combined stage may have a first crossover frequency k y For frequencies above the first crossover frequency, the combined stage will use the spectral components of the five waveform-encoded signals 210a-e. For frequencies above the first crossover frequency, the combined stage will use the upmix signal 404 interleaved with a further waveform-encoded signal 710.

[0082] The interleaving stage 714 may operate under the control of a control signal. To this end, the decoder 100′ may receive, for example through the third receiving stage 616, a control signal indicating how to interleave the further waveform-encoded signal with one of the M upmix signals. For example, the control signal may indicate a frequency range and a time range in which the further waveform-encoded signal 710 should be interleaved with one of the upmix signals 404. For example, the frequency range and the time range may be expressed in terms of time / frequency tiles in which the interleaving should be performed. The time / frequency tiles may be time / frequency tiles in terms of a time / frequency grid in the QMF domain in which the interleaving is performed.

[0083] The control signal may use a vector, such as a binary vector, to indicate the time / frequency tiles on which interleaving should be performed. Specifically, there may be a first vector for frequency indication, indicating the frequencies on which interleaving should be performed. The indication may be made, for example, by indicating a logical 1 for the corresponding frequency interval in the first vector. There may also be a second vector for time indication, indicating the time interval on which interleaving should be performed. The indication may be made, for example, by indicating a logical 1 for the corresponding time interval in the second vector. For this purpose, a time frame is generally divided into multiple time slots so that the time indication may be made on a subframe basis. A time / frequency matrix may be constructed by intersecting the first and second vectors. For example, the time / frequency matrix may be a binary matrix containing a logical 1 for each time / frequency tile on which the first and second vectors indicate a logical 1. The interleaving stage 714 may then use the time / frequency matrix to perform interleaving, such that one or more of the upmix signals 404 are replaced by the further waveform-encoded signal 710, for example, for time / frequency tiles indicated by, for example, a logical 1 in the time / frequency matrix.

[0084] It is noted that the vector may use other schemes rather than a binary scheme to indicate the time / frequency tiles on which interleaving should be performed. For example, the vector might use a first value, such as zero, to indicate that interleaving should not be performed, and a second value to indicate that interleaving should be performed for the particular channel identified by the second value.

[0085] FIG. 5 shows, by way of example, a generalized block diagram of an encoding system 500 suitable for a multi-channel audio processing system for encoding M channels, according to one embodiment.

[0086] In the exemplary embodiment illustrated in FIG. 5, encoding of 5.1 surround sound is described. Therefore, in the illustrated example, M is set to 5. It may be noted that a low-frequency effect signal is not mentioned in the illustrated embodiment or in the drawings. This does not mean that any low-frequency effects are ignored. Low-frequency effects (Lfe) are added to the bitstream 552 in any appropriate manner well known by those skilled in the art. It may also be noted that the described encoder is equally well suited for encoding other types of surround sound, such as 7.1 or 9.1 surround sound. In the encoder 500, five signals 502, 504 are received at a receiving stage (not shown). The encoder 500 includes a first waveform encoding stage 506 configured to receive the five signals 502, 504 from the receiving stage and generate five waveform-encoded signals 518 by waveform-encoding the five signals 502, 504 individually. The waveform coding stage 506 may, for example, perform an MDCT transform on each of the five received signals 502, 504. As discussed with respect to the decoder, the encoder may choose to code each of the five received signals 502, 504 using an MDCT transform with independent windowing. This may allow for improved coding quality and, therefore, increased quality of the decoded signals.

[0087] The five waveform-coded signals 518 are waveform-coded for a frequency range corresponding to frequencies up to a first crossover frequency. Thus, the five waveform-coded signals 518 include spectral coefficients corresponding to frequencies up to the first crossover frequency, which can be obtained by low-pass filtering each of the five waveform-coded signals 518. The five waveform-coded signals 518 are then quantized 520 according to a psychoacoustic model. The psychoacoustic model is configured to reproduce as accurately as possible the coded signals as perceived by a listener when decoded at the decoder side of the system, taking into account the bit rate available in the multi-channel audio processing system.

[0088] As discussed above, the encoder 500 performs hybrid coding, including discrete multi-channel coding and parametric coding. Discrete multi-channel coding, as described above, is performed on each of the input signals 502, 504 in the waveform coding stage 506 for frequencies up to the first crossover frequency. Parametric coding is performed on the decoder side for frequencies above the first crossover frequency so that the five input signals 502, 504 can be reconstructed from the N downmix signals. In the illustrated example in FIG. 5 , N is set to 2. Downmixing of the five input signals 502, 504 is performed in the downmixing stage 534. The downmixing stage 534 advantageously operates in the QMF domain. Therefore, before being input to the downmixing stage 534, the five signals 502, 504 are converted to the QMF domain by the QMF analysis stage 526. The downmixing stage performs a linear downmixing operation on the five signals 502, 504 and outputs two downmix signals 544, 546.

[0089] These two downmix signals 544, 546 are received by the second waveform coding stage 508 after being transformed back into the time domain by an inverse QMF transform 554. The second waveform coding stage 508 generates two waveform-coded downmix signals by waveform coding the two downmix signals 544, 546 for a frequency range corresponding to frequencies between the first and second crossover frequencies. The waveform coding stage 508 may, for example, perform an MDCT transform on each of the two downmix signals. Thus, the two waveform-coded downmix signals include spectral coefficients corresponding to frequencies between the first and second crossover frequencies. The two waveform-coded downmix signals are then quantized 522 according to a psychoacoustic model.

[0090] To enable the decoder side to restore frequencies above the second crossover frequency, high frequency restoration (HFR) parameters 538 are extracted from the two downmix signals 544, 546. These parameters are extracted in the HFR encoding stage 532.

[0091] To enable the decoder side to reconstruct five signals from the two downmix signals 544, 546, the five input signals 502, 504 are received by a parametric encoding stage 530. The five signals 502, 504 are parametrically encoded for a frequency range corresponding to frequencies above a first crossover frequency. The parametric encoding stage 530 is then configured to extract upmix parameters 536 that enable upmixing of the two downmix signals 544, 546 into five reconstructed signals corresponding to the five input signals 502, 504 (i.e., five channels in the encoded 5.1 surround sound) for a frequency range above the first crossover frequency. It may be noted that the upmix parameters 536 are extracted only for frequencies above the first crossover frequency. This may reduce the complexity of the parametric encoding stage 530 and the bit rate of the corresponding parametric data.

[0092] It may be noted that downmixing 534 can be achieved in the time domain. In such a case, since HFR encoding stage 532 generally operates in the QMF domain, QMF analysis stage 526 should be located downstream of downmixing stage 534 and before HFR encoding stage 532. In this case, inverse QMF stage 554 can be omitted.

[0093] The encoder 500 further comprises a bitstream generation stage, i.e., a bitstream multiplexer 524. According to an exemplary embodiment of the encoder 500, the bitstream generation stage is configured to receive five coded and quantized signals 548, two parameter signals 536, 538, and two coded and quantized downmix signals 550, which are converted by the bitstream generation stage 524 into a bitstream 552 for further distribution in a multi-channel audio system.

[0094] In the described multi-channel audio system, a maximum available bit rate often exists, for example, when streaming audio over the Internet. Because the characteristics of each time frame of the input signals 502, 504 are different, the exact same allocation of bits may not be used between the five waveform-encoded signals 548 and the two downmix waveform-encoded signals 550. Furthermore, each individual signal 548, 550 may require more or fewer allocated bits so that the signal can be restored according to a psychoacoustic model. According to an exemplary embodiment, the first and second waveform-encoding stages 506, 508 share a common bit reservoir. The available bits per encoded frame are initially distributed between the first and second waveform-encoding stages 506, 508 depending on the characteristics of the signal to be encoded and the current psychoacoustic model. As explained above, the bits are then distributed between the individual signals 548, 550. The number of bits used for the high-frequency restoration parameters 538 and the upmix parameters 536 is, of course, taken into account when distributing the available bits. Care is taken to adjust the psychoacoustic models for the first and second waveform coding stages 506, 508 with respect to the number of bits allocated in a particular time frame for a perceptually smooth transition around the first crossover frequency.

[0095] Figure 8 illustrates an alternative embodiment of an encoding system 800. The difference between the encoding system 800 of Figure 8 and the encoding system 500 of Figure 5 is that the encoder 800 is arranged to generate a further waveform-coded signal by waveform encoding one or more of the input signals 502, 504 for a frequency range corresponding to a subset of the frequency range above the first crossover frequency.

[0096] To this end, the encoder 800 includes an interleave detection stage 802. The interleave detection stage 802 is configured to identify portions of the input signals 502, 504 that are not well reconstructed by the parametric reconstructions encoded by the parametric encoding stage 530 and the high-frequency reconstruction encoding stage 532. For example, the interleave detection stage 802 may compare the input signals 502, 504 to the parametric reconstructions of the input signals 502, 504 defined by the parametric encoding stage 530 and the high-frequency reconstruction encoding stage 532. Based on the comparison, the interleave detection stage 802 may identify a subset 804 of the frequency range above the first crossover frequency to be waveform coded. The interleave detection stage 802 may similarly identify a time range over which the identified subset 804 of the frequency range above the first crossover frequency should be waveform coded. The identified frequency and time subsets 804, 806 may be input to the first waveform encoding stage 506. Based on the received frequency and time subsets 804, 806, the first waveform encoding stage 506 generates a further waveform-encoded signal 808 by waveform encoding one or more of the input signals 502, 504 for the time and frequency ranges identified by the subsets 804, 806. The further waveform-encoded signal 808 may then be coded and quantized by stage 520 and added to a bitstream 846.

[0097] The interleave detection stage 802 may further include a control signal generation stage configured to generate a control signal 810 indicating how the further waveform-encoded signal is to be interleaved with a parametric reconstruction of one of the input signals 502, 504 in the decoder. As described with reference to Figure 7, for example, the control signal may indicate a frequency range and a time range in which the further waveform-encoded signal should be interleaved with the parametric reconstruction. The control signal may be added to the bitstream 846.

[0098] "Equivalents, Extensions, Substitutes, and Others" Further embodiments of the present disclosure will be apparent to those skilled in the art after considering the above description. Although the description and drawings disclose embodiments and examples, the present disclosure is not limited to these particular examples. Many modifications and variations can be made without departing from the scope of the present disclosure, which is defined by the appended claims. Any reference signs appearing in the claims should not be construed as limiting their scope.

[0099] Furthermore, variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the present disclosure, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude a plurality. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

[0100] The systems and methods disclosed above may be implemented as software, firmware, hardware, or a combination thereof. In hardware implementations, the division of tasks among functional units referred to in the above description does not necessarily correspond to a division into physical units; conversely, one physical component may have multiple functions, and one task may be performed by several cooperating physical components. Certain or all components may be implemented as software executed by a digital signal processor or microprocessor, or as hardware or an application-specific integrated circuit. Such software may be distributed via computer-readable media, which may include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those skilled in the art, the term computer storage media includes both volatile and non-volatile, removable and non-removable media, implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and that can be accessed by a computer. Additionally, those skilled in the art will appreciate that communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport means, and includes any information delivery media.

Claims

1. 1. A method of decoding a time frame of an encoded audio bitstream in an audio processing system, the method comprising: extracting from the encoded audio bitstream a first waveform-coded signal including spectral coefficients corresponding to frequencies up to a first crossover frequency for a time frame; performing parametric decoding above a second crossover frequency in a reconstruction range for the time frame to generate a reconstructed signal, the second crossover frequency being above the first crossover frequency, and the parametric decoding using reconstruction parameters derived from the encoded audio bitstream to generate the reconstructed signal; extracting from the encoded audio bitstream a second waveform-encoded signal including spectral coefficients corresponding to a subset of frequencies above the first crossover frequency for the time frame, the first waveform-encoded signal and the second waveform-encoded signal being signals representing a waveform of an audio signal in the frequency domain; interleaving the second waveform-encoded signal with the reconstructed signal to generate an interleaved signal for the time frame, the interleaving comprising: (i) adding the second waveform-encoded signal with the reconstructed signal; (ii) combining the second waveform-encoded signal with the reconstructed signal; or (iii) replacing the reconstructed signal with the second waveform-encoded signal; A method comprising:

2. 10. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of claim 1.

3. 1. An audio decoder for decoding a time frame of an encoded audio bitstream, the audio decoder comprising: a first demultiplexer for extracting from the encoded audio bitstream a first waveform-encoded signal including spectral coefficients corresponding to frequencies up to a first crossover frequency for a time frame; a parametric decoder that performs parametric decoding above a second crossover frequency in a reconstruction range for the time frame to generate a reconstructed signal, the second crossover frequency being above the first crossover frequency, and the parametric decoding uses reconstruction parameters derived from the encoded audio bitstream to generate the reconstructed signal; a second demultiplexer for extracting from the encoded audio bitstream a second waveform-encoded signal including spectral coefficients corresponding to a subset of frequencies above the first crossover frequency for the time frame, the first waveform-encoded signal and the second waveform-encoded signal being signals representing a waveform of an audio signal in the frequency domain; and an interleaver for interleaving the second waveform-coded signal with the reconstructed signal to generate an interleaved signal for the time frame, the interleaving including (i) adding the second waveform-coded signal with the reconstructed signal, (ii) combining the second waveform-coded signal with the reconstructed signal, or (iii) replacing the reconstructed signal with the second waveform-coded signal; and 1. An audio decoder comprising:

Citation Information

Patent Citations

  • Multi-channel / cue coding / decoding of audio signal

    JP2004078183A

  • Apparatus and method for encoding / decoding using phase information and residual signal

    JP2013508770A

  • JPP7413418B