Apparatus for encoding or decoding an encoded multi-channel signal using a supplemental signal generated by a wideband filter - Patent Application 20070122997

The hybrid approach using a decorrelation filter and multi-channel processor for multi-channel audio decoding enhances audio quality and efficiency by generating filler signals with appropriate temporal and spectral characteristics, addressing issues in existing codecs like xHE-AAC and AMR-WB+.

JP7804634B2Active Publication Date: 2026-01-22FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023206539
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-07-28
Filing Date
2023-12-07
Publication Date
2026-01-22
Estimated Expiration
2038-07-26

AI Technical Summary

Technical Problem

Existing multi-channel audio codecs like xHE-AAC and AMR-WB+ face issues with speech signals, producing unnatural tones and artifacts due to inconsistent output for different sampling rates, and inefficiencies in bit rate usage, especially for reverberant or out-of-phase signals.

Method used

A hybrid approach using a decorrelation filter generates filler signals in the time domain, which are then processed in the spectral domain by a multi-channel processor to create a decoded multi-channel signal, maintaining high audio quality while minimizing bit rate.

Benefits of technology

This method improves audio quality for speech signals by generating filler signals with temporal and spectral characteristics closer to the input, addressing inconsistencies across sampling rates and enhancing stereo image stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007804634000107
    Figure 0007804634000107
  • Figure 0007804634000108
    Figure 0007804634000108
  • Figure 0007804634000109
    Figure 0007804634000109
Patent Text Reader

Abstract

To provide a device and method for combining frequency domain multichannel processing with time domain uncorrelating to obtain decoded multichannel signals with high audio quality.SOLUTION: A device for decoding a multichannel signal includes: a base channel decoder 700 for decoding an encoded base channel to obtain a decoded base channel; an uncorrelating filter 800 for filtering at least a portion of the decoded base channel to obtain a supplemental signal; and a multichannel processor 900 for performing multichannel processing using spectral representation of the decoded base channel and spectral representation of the supplemental signal. The uncorrelating filter is a wideband filter, and the multichannel processor applies narrowband processing to the spectral representation of the decoded base channel and the spectral representation of the supplemental signal.SELECTED DRAWING: Figure 7a
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to audio processing, and in particular to multi-channel audio processing in an apparatus or method for decoding an encoded multi-channel signal. [Background technology]

[0002] The current state-of-the-art codec for parametric coding of stereo signals at low bit rates is the MPEG codec xHE-AAC. It features a fully parametric stereo coding aspect based on a mono downmix and stereo parameters evaluated in the subbands, namely the inter-channel level difference (ILD) and the inter-channel coherence (ICC). The output is synthesized from the mono downmix by matrixing, in each subband, the subband downmix signal with a decorrelated version of that subband downmix signal obtained by applying the subband filters in a QMF filterbank.

[0003] There are several drawbacks associated with xHE-AAC for coding speech items. The filter that generates the synthetic second signal generates a highly reverberant version of the input signal, which requires a ducker. Therefore, the processing significantly distorts the spectral shape of the input signal over time. While this works well for many types of signals, speech signals with rapidly changing spectral envelopes can produce unnatural tones and audible artifacts such as double talk or ghost voices. Furthermore, the filter depends on the time resolution of the underlying QMF filter bank, which varies with the sampling rate. Therefore, the output signal is inconsistent for different sampling rates.

[0004] Separately, the 3GPP (registered trademark) codec AMR-WB+ features a semi-parametric stereo mode supporting bit rates from 7 to 48 kbit / s. It is based on a mid / side transformation of the left and right input channels. In the low frequency range, the side signal s is predicted by the mid signal m to obtain a balance gain, and both m and the prediction residual are encoded and transmitted to the decoder along with the prediction coefficients. In the mid frequency range, only the downmix signal m is coded, and the missing signal s is predicted from m using a low-order FIR filter calculated in the encoder. This is combined with bandwidth expansion for both channels. While this codec typically produces a more natural sound than xHE-AAC for speech, it faces several problems. The procedure of predicting s by m using a low-order FIR filter does not work very well when the input channels are only weakly correlated, such as in the case of reverberant speech signals or double talk. Also, this codec cannot handle out-of-phase signals, which can lead to a significant loss of quality, and the stereo image of the decoded output usually feels highly compressed. Furthermore, the method is not fully parametric, and therefore not efficient in terms of bit rate.

[0005] In general, fully parametric methods may result in a degradation of audio quality, since the signal parts lost due to parametric encoding are not reconstructed at the decoder side.

[0006] On the other hand, waveform-preserving procedures such as mid / side coding do not offer the significant bit-rate savings that can be obtained from parametric multi-channel coders. Summary of the Invention [Problem to be solved by the invention]

[0007] It is an object of the present invention to provide an improved concept for decoding an encoded multi-channel signal. [Means for solving the problem]

[0008] This object is achieved by an apparatus for decoding an encoded multi-channel signal, a method for decoding an encoded multi-channel signal according to claim 37, a computer program according to claim 38, and an audio signal decorrelator according to claim 39, a method for decorrelating an audio input signal according to claim 49 or a computer program according to claim 50.

[0009] The present invention is based on the discovery that a hybrid approach is useful for decoding encoded multi-channel signals. This hybrid approach relies on the use of filler signals generated by a decorrelation filter, which are then used by a multi-channel processor, such as a parametric or other multi-channel processor, to generate a decoded multi-channel signal. In particular, the decorrelation filter is a wideband filter, and the multi-channel processor is configured to apply narrowband processing to the spectral representation. Thus, the filler signals are preferably generated in the time domain, for example by an all-pass filter procedure, and multi-channel processing is performed in the spectral domain using spectral representations of the decoded base channels and further using spectral representations of the filler signals generated from the filler signals calculated in the time domain.

[0010] Thus, the advantages of frequency-domain multi-channel processing on the one hand and time-domain decorrelation on the other hand are combined in a useful way to obtain a decoded multi-channel signal with high audio quality. Nevertheless, the bit rate for transmitting the encoded multi-channel signal is kept as low as possible due to the fact that the encoded multi-channel signal is typically not in a waveform-preserving encoding format but in a parametric multi-channel coding format, for example. Therefore, to generate the supplementary signal, only data available at the decoder, such as the decoded base channel, is used, and in certain embodiments, additional stereo parameters, such as gain parameters or prediction parameters, or ILD, ICC, or any other stereo parameters known in the art, are used.

[0011] Next, we will describe some preferred embodiments. The most efficient way to code a stereo signal is to use parametric methods such as binaural cue coding or parametric stereo. These methods aim to recreate a spatial impression from a mono downmix by restoring some spatial cues within the subbands and are themselves based on psychoacoustics. Another approach that relies on parametric methods is to exploit inter-channel redundancy and attempt to parametrically model one channel with another. While this approach can recover parts of the secondary channel from the primary channel, residual components usually remain. Omitting these components usually leads to an unstable stereo image in the decoded output. Therefore, it is necessary to supplement these residual components with appropriate replacements. Because such replacements are blind, it is safest to obtain these parts from a second signal with similar temporal and spectral characteristics to the downmix signal.

[0012] Therefore, embodiments of the present invention are particularly useful in the context of parametric audio coders, and in particular parametric audio decoders, where a replacement for the missing residual part is extracted from an artificial signal generated by a decoder-side decorrelation filter.

[0013] Further embodiments relate to procedures for generating an artificial signal. Some embodiments relate to a method for generating an artificial second channel from which a replacement for the missing residual part is extracted, and its use in a fully parametric stereo coder called extended stereo filling. This signal is more suitable for coding speech signals than an xHE-AAC signal because its spectral shape is temporally closer to the input signal. Because it is generated in the time domain by applying a special filter structure, it is independent of the filter bank in which the stereo upmix is ​​performed. Therefore, it can be used in various upmix procedures. For example, in xHE-AAC, it can be used to replace the artificial signal after transformation to the QMF domain, which is thought to improve performance for speech, and even in the mid-band of AMR-WB+, it can be used to replace the residual in mid / side prediction, which is thought to improve performance for weakly correlated input channels and improve the stereo image. This is particularly interesting for codecs with different stereo modes (such as time-domain and frequency-domain stereo processing).

[0014] In a preferred embodiment, the decorrelation filter comprises at least one all-pass filter cell comprising two Schrader all-pass filter cells nested in a third Schrader all-pass filter, and / or the all-pass filter comprises at least one all-pass filter cell comprising two cascaded Schrader all-pass filters, the input to the first cascaded Schrader all-pass filter and the output from the second cascaded Schrader all-pass filter being connected in the signal flow direction before the delay stage of the third Schrader all-pass filter.

[0015] In a further embodiment, several such all-pass filter cells, including three nested Schroeder all-pass filters, are cascaded to obtain a particularly useful all-pass filter with a good impulse response for the purposes of stereo or multi-channel decoding.

[0016] While some aspects of the present invention are discussed herein with respect to stereo decoding, which generates a left upmix channel and a right upmix channel from a mono base channel, it should be emphasized that the present invention is also applicable to multi-channel decoding, in which, for example, a four-channel signal is encoded using two base channels, the first two upmix channels are generated from the first base channel, and the third and fourth upmix channels are generated from the second base channel. In another alternative, the present invention is also useful for generating three or more upmix channels from a single base channel, preferably always using the same filler signals. However, in all such procedures, the filler signals are generated in a wideband manner, i.e., preferably in the time domain, and the multi-channel processing for generating two or more upmix channels from the decoded base channel is performed in the frequency domain.

[0017] The decorrelation filter preferably operates entirely in the time domain. However, other hybrid approaches are also useful, for example, where decorrelation is performed by decorrelating the low-band portion on the one hand and the high-band portion on the other hand, while the multi-channel processing is performed with a much higher spectral resolution. Thus, for example, the spectral resolution of the multi-channel processing may be as high as the individual processing of each DFT or FFT line, for example, where parametric data is provided for several bands, each band including, for example, two, three, or even more DFT / FFT / MDCT lines, and the filtering of the decoded base channel to obtain the supplemental signal is performed in a wideband manner, i.e., in the time domain, or in a semi-wideband manner, such as in the low-band and the high-band, or in three different bands. Therefore, in either case, the spectral resolution of the stereo processing typically performed on individual line or subband signals is the highest spectral resolution. Typically, the stereo parameters generated and transmitted in the encoder and used by the preferred decoder have a medium spectral resolution. Thus, parameters are given for bands, which can have different bandwidths, but each band includes at least two or more line or subband signals that are generated and used by the multi-channel processor. Furthermore, the spectral resolution of the decorrelation filtering is very low, either extremely low in the case of time-domain filtering, or medium in the case of generating different decorrelated signals for different bands, but this medium spectral resolution is still lower than the resolution at which parameters for parametric processing are given.

[0018] In a preferred embodiment, the filter characteristic of the decorrelation filter is an all-pass filter having a constant size area throughout the spectral range of interest. However, other decorrelation filters that do not have this ideal all-pass filter behavior are also useful in preferred embodiments, provided that the area over which the filter characteristic is constant size is greater than the spectral granularity of the spectral representation of the decoded base channel and the spectral granularity of the spectral representation of the filler signal.

[0019] In this way, it is ensured that the spectral granularity of the filler signals or the decoded base channels on which multi-channel processing is performed does not affect the decorrelation filtering process, so that high-quality filler signals are generated, preferably adjusted using energy normalization factors, and then used to generate two or more upmix channels.

[0020] Furthermore, it should be noted that the generation of decorrelated signals such as those described with respect to Figures 4, 5 or 6 discussed below can be used in the context of a multi-channel decoder, but also in any other application in which decorrelated signals are useful, such as, for example, rendering audio signals, reverberation operations, etc.

[0021] Preferred embodiments will now be described with reference to the accompanying drawings. [Brief explanation of the drawings]

[0022] [Figure 1a] 1 shows artificial signal generation when used with the EVS core coder. [Figure 1b] 10 illustrates artificial signal generation when used with an EVS core coder according to another embodiment. [Figure 2a] Integration into DFT stereo processing including time-domain bandwidth extension upmix is ​​shown. [Figure 2b] 1 illustrates an integration into DFT stereo processing with a time-domain bandwidth extension upmix according to another embodiment. [Figure 3] Illustrates integration into a system with multiple stereo processing units. [Figure 4] A basic all-pass unit is shown. [Figure 5] 1 shows an all-pass filter unit. [Figure 6] 1 shows the impulse response of a preferred all-pass filter. [Figure 7a] 1 shows an apparatus for decoding an encoded multi-channel signal. [Figure 7b] 1 shows a preferred embodiment of the decorrelation filter. [Figure 7c] 1 shows a combination of a base channel decoder and a spectral transformer. [Figure 8] 1 shows a preferred embodiment of a multi-channel processor. [Figure 9a] 1 shows a further embodiment of an apparatus for decoding a multi-channel signal encoded using bandwidth extension processing; [Figure 9b] 1 illustrates a preferred embodiment for generating the compressed energy normalization coefficients. [Figure 10] 10 shows an apparatus for decoding an encoded multi-channel signal according to a further embodiment that operates using a channel transform in a base channel decoder. [Figure 11] 1 shows the cooperation between the resampler of the base channel decoder and the subsequent decorrelation filter. [Figure 12] 1 shows an exemplary parametric multi-channel encoder useful in an apparatus for decoding according to the present invention; [Figure 13] 1 shows a preferred embodiment of an apparatus for decoding an encoded multi-channel signal. [Figure 14] 2 shows a further preferred embodiment of a multi-channel processor. DETAILED DESCRIPTION OF THE INVENTION

[0023] 7a shows a preferred embodiment of an apparatus for decoding an encoded multi-channel signal, wherein the encoded multi-channel signal includes an encoded base channel, which is input to a base channel decoder 700 for decoding the encoded base channel to obtain a decoded base channel.

[0024] Additionally, the decoded base channel is input to a decorrelation filter 800 for filtering at least a portion of the decoded base channel to obtain a filler signal.

[0025] Both the decoded base channel and the filler signal are input to a multi-channel processor 900 for performing multi-channel processing using a spectral representation of the decoded base channel and also using a spectral representation of the filler signal. The multi-channel processor outputs a decoded multi-channel signal including, for example, a left upmix channel and a right upmix channel in a stereo processing situation, or three or more upmix channels in the case of multi-channel processing covering three or more output channels.

[0026] The decorrelation filter 800 is configured as a wideband filter, and the multi-channel processor 900 is configured to apply narrowband processing to the spectral representations of the decoded base channel and the filler signal. Importantly, wideband filtering also occurs if the signal to be filtered has been downsampled from a high sampling rate, such as 22 kHz or less, to 16 kHz or 12.8 kHz.

[0027] The multi-channel processor therefore operates at a spectral granularity significantly higher than the spectral granularity of the generation of the filler signals, in other words, the filter characteristics of the decorrelation filter are selected such that the area of ​​the filter characteristic of a certain magnitude is larger than the spectral granularity of the spectral representation of the decoded base channel and the spectral granularity of the spectral representation of the filler signals.

[0028] Thus, for example, if the spectral granularity of the multi-channel processor is such that the upmixing process is performed for each spectral line of a 1024-line DFT spectrum, the decorrelation filter is defined in such a way that the region of the decorrelation filter's constant magnitude filter characteristic has a frequency width higher than two or more spectral lines of the DFT spectrum. Typically, the decorrelation filter operates in the time domain, and a spectral band of, for example, 20 Hz to 20 kHz is used. Such filters are known as all-pass filters. It should be noted that although a perfectly constant magnitude range is typically not obtainable with an all-pass filter, variations from a constant magnitude by + / - 10% of the mean value are also considered useful for all-pass filters and therefore correspond to a "constant magnitude filter characteristic."

[0029] 7b shows an implementation of a decorrelation filter 800 with a time-domain filter stage 802 followed by a spectral transform 804 that produces a spectral representation of the concatenated supplemental signal. The spectral transformer 804 is typically implemented as an FFT or DFT processor, although other time-frequency domain transform algorithms are also useful.

[0030] FIG. 7c shows a preferred embodiment of the cooperation between the base channel decoder 700 and the base channel spectral converter 902. Typically, the base channel decoder is configured to operate as a time-domain base channel decoder that generates a time-domain base channel signal, while the multi-channel processor 900 operates in the spectral domain. Thus, the multi-channel processor 900 of FIG. 7a has the base channel spectral converter 902 of FIG. 7c as an input stage, so that the spectral representation of the base channel spectral converter 902 is transferred to the multi-channel processor processing elements shown, for example, in FIG. 8, 13, 14, 9a, or 10. In this context, it should be generally noted that reference numbers beginning with "7" represent elements that preferably belong to the base channel decoder 700 of FIG. 7a. Elements with reference numbers beginning with "8" preferably belong to the decorrelation filter 800 of FIG. 7a, and elements with reference numbers beginning with "9" in the figures preferably belong to the multi-channel processor 900 of FIG. 7a. However, it should be noted that the separation between individual elements is done here merely for the purpose of explaining the invention, and that actual implementations may have different processing blocks, typically hardware processing blocks, or software processing blocks, or mixed hardware / software processing blocks, separated in a manner different from the logical separation shown in Figure 7a and other figures.

[0031] FIG. 4 illustrates a preferred embodiment of filter stage 802, designated as 802′. In particular, FIG. 4 illustrates a basic all-pass unit that may be included in a decorrelation filter, which may be included in the decorrelation filter alone or together with more such cascaded all-pass units, as shown, for example, in FIG. 5. While FIG. 5 illustrates decorrelation filter 802 having typically five cascaded basic all-pass units 502, 504, 506, 508, 510, each of the basic all-pass units may be implemented as outlined in FIG. 4. However, the decorrelation filter may alternatively include a single basic all-pass unit 403 of FIG. 4, thus representing an alternative implementation of decorrelation filter stage 802′.

[0032] Preferably, each elementary all-pass unit comprises two Schrader all-pass filters 401, 402 nested within a third Schrader all-pass filter 403. In this embodiment, the all-pass filter cell 403 is connected to two cascaded Schrader all-pass filters 401, 402, with the input to the first cascaded Schrader all-pass filter 401 and the output from the second cascaded Schrader all-pass filter 402 being connected, in the direction of signal flow, before the delay stage 423 of the third Schrader all-pass filter.

[0033] In particular, the all-pass filter shown in FIG. 4 comprises a first adder 411, a second adder 412, a third adder 413, a fourth adder 414, a fifth adder 415, and a sixth adder 416, a first delay stage 421, a second delay stage 422, and a third delay stage 423, a first forward feed 431 having a first forward gain, a first reverse feed 441 having a first reverse gain, a second forward feed 442 having a second forward gain, and a second reverse feed 432 having a second reverse gain, as well as a third forward feed 443 having a third forward gain and a third reverse feed 433 having a third reverse gain.

[0034] 4 is as follows: the input to the first adder 411 corresponds to the input to the all-pass filter 802; the second input to the first adder 411 is connected to the output of the third filter delay stage 423 and includes a third backward feed 433 having a third backward gain; the output of the first adder 411 is connected to the input of the second adder 412 and is connected to the input of the sixth adder 416 via a third forward feed 443 having a third forward gain; the input to the second adder 412 is connected to the first delay stage 421 via a first backward feed 441 having a first backward gain; the output of the second adder 412 is connected to the input of the first delay stage 421 and is connected to the input of the third adder 413 via a first forward feed 431 having a first forward gain. The output of the first delay stage 421 is connected to a further input of a third adder 413. The output of the third adder 413 is connected to an input of a fourth adder 414. A further input to the fourth adder 414 is connected to the output of the second delay stage 422 via a second backward feed 432 having a second backward gain. The output of the fourth adder 414 is connected to an input to the second delay stage 422 and to an input to the fifth adder 415 via a second forward feed 442 having a second forward gain. The output of the second delay stage 421 is connected to a further input to the fifth adder 415. The output of the fifth adder 415 is connected to an input of the third delay stage 423. The output of the third delay stage 423 is connected to an input to a sixth adder 416. A further input to the sixth adder 416 is connected to the output of the first adder 411 via a third forward feed 443 having a third forward gain. The output of the sixth adder 416 corresponds to the output of the all-pass filter 802.

[0035] Preferably, as shown in FIG. 8, the multi-channel processor 900 is configured to determine the first and second upmix channels using different weighted combinations of the spectral bands of the decoded base channel and the corresponding spectral bands of the filler signal. In particular, the different weighted combinations depend on prediction coefficients and / or gain coefficients derived from encoded parametric information included in the encoded multi-channel signal. Furthermore, the weighted combinations preferably depend on envelope normalization coefficients or, preferably, on energy normalization coefficients calculated using the spectral bands of the decoded base channel and the corresponding spectral bands of the filler signal. Thus, the processor 904 of FIG. 8 receives the spectral representations of the decoded base channel and the spectral representations of the filler signal and outputs the first and second upmix channels, preferably in the time domain, where the prediction coefficients, gain coefficients, and energy normalization coefficients are input in a band-by-band manner, where these coefficients are used for all spectral lines within a band but vary for different bands, and where this data is obtained from the encoded signal or determined locally in the decoder.

[0036] In particular, the prediction coefficients and gain coefficients typically correspond to encoded parameters that are decoded at the decoder side and then used for parametric stereo upmixing. Conversely, the energy normalization coefficients are typically calculated at the decoder side using the spectral bands of the decoded base channel and the spectral bands of the supplemental signal. The same applies to the envelope normalization coefficients. Preferably, the envelope normalization corresponds to band-by-band energy normalization.

[0037] Although the present invention is described with reference to the specific reference encoder shown in Figure 12 and the specific decoder shown in Figure 13 or 14, it should be noted that the generation and application of wideband filler signals in multi-channel stereo decoding operating in a narrowband spectral domain is also applicable to any other parametric stereo encoding techniques known in the art, such as the HE-AAC standard, or the MPEG Surround standard, or binaural cue coding (BCC coding), or any other stereo encoding / decoding tool, or any other multi-channel encoding / decoding tool.

[0038] 9a shows a further preferred embodiment of a multi-channel decoder comprising a multi-channel processor stage 904 that generates a first upmix channel and a second upmix channel, and subsequent time-domain bandwidth extension elements 908, 910 that perform time-domain bandwidth extension on the first upmix channel and the second upmix channel separately in a guided or unguided manner. Typically, a windower and energy normalization coefficient calculator 912 is provided to calculate an energy normalization coefficient used by the multi-channel processor 904. However, in alternative embodiments described with respect to FIGS. 1a or 1b and 2a or 2b, the bandwidth extension is performed on the mono or decoded core signal, and only the single stereo processing element 960 of FIG. 2a or 2b is provided to generate from the high-band mono signal a high-band left channel signal and a high-band right channel signal that are subsequently added to the low-band left channel signal and the low-band right channel signal using summers 994a and 994b.

[0039] This addition shown in Figure 2a or 2b can be performed, for example, in the time domain. Thus, block 960 generates a time-domain signal. This is a preferred embodiment. However, alternatively, the stereo processing 904 of Figure 2a or 2b and the left and right channel signals from block 960 can be generated in the spectral domain, with adders 994a and 994b being realized, for example, by a synthesis filter bank, with the low-band data from block 904 input to the low-band input of the synthesis filter bank, the high-band output of block 960 input to the high-band input of the synthesis filter bank, and the output of the synthesis filter bank being the corresponding left channel time-domain signal or right channel time-domain signal.

[0040] Preferably, the windower and coefficient calculator 912 of FIG. 9a generates and calculates energy values ​​for the highband signal, e.g. as also shown in 961 of FIG. 1a or 1b, and uses this energy estimate to generate the first and second upmix channels for the highband, as described below in a preferred embodiment with respect to equations 28-31.

[0041] Preferably, the processor 904 for calculating the weighted combinations receives as input an energy normalization coefficient for each band. However, in a preferred embodiment, compression of the energy normalization coefficients is performed, and the compressed energy normalization coefficients are used to calculate different weighted combinations. Thus, with respect to FIG. 8, the processor 904 receives compressed energy normalization coefficients instead of uncompressed energy normalization coefficients. This procedure is illustrated in FIG. 9b for a different embodiment. Block 920 receives the energy of the residual or filler signal for each time / frequency bin and the energy of the decoded base channel for each time and frequency bin, and then calculates absolute energy normalization coefficients for bands containing several such time / frequency bins. Next, in block 921, compression of the energy normalization coefficients is performed, which may be, for example, using a logarithmic function as discussed below with respect to Equation 22.

[0042] Based on the compressed energy normalization coefficients generated by block 921, different procedures are provided for generating compressed energy normalization coefficients. In a first alternative, a function is applied to the compressed coefficients, as shown in 922, which is preferably a nonlinear function. Next, in block 923, the evaluated coefficients are expanded to obtain specific compressed energy normalization coefficients. Thus, block 922 can be implemented, for example, in the functional expression of equation (22) described below, and block 923 is performed by the "exponential" function in equation (22). However, another alternative that results in similar compressed energy normalization coefficients is shown in blocks 924 and 925. In block 924, an evaluation factor is determined, and in block 925, the evaluation factor is applied to the energy normalization coefficients obtained from block 920. Thus, the application of a factor to the energy normalization coefficients as schematically shown in block 912 can be implemented, for example, by equation (27) described below.

[0043] Thus, for example, as shown below in Equation 27, an evaluation factor is determined that is simply the energy normalization factor as determined by block 920 without actually performing any special function evaluation. TIFF0007804634000001.tif412. Therefore, the calculation of block 925 can also be omitted, that is, the specific calculation of the compressed energy normalization coefficient is not necessary once the original uncompressed energy normalization coefficient and further operands in the multiplication, such as the evaluation coefficient and the spectral value of the fill signal, are multiplied together to obtain the normalized fill signal spectral line.

[0044] 10 shows a further embodiment in which the encoded multi-channel signal is not just a mono signal but also includes, for example, an encoded intermediate signal and an encoded side signal. In such a situation, the base channel decoder 700 not only decodes the encoded intermediate signal and the encoded side signal, or generally the encoded first signal and the encoded second signal, but also performs a channel transform 705 in the form of, for example, a mid / side transform and an inverse mid / side transform to calculate a primary channel such as L and a secondary channel such as R, or the transform is a Karhunen Loeve transform.

[0045] However, the result of the channel transformation, particularly the decoding operation, is that the primary channel is a wideband channel while the secondary channel is a narrowband channel. Next, the wideband channel is input to the decorrelation filter 800, and high-pass filtering is performed in block 930 to generate a de-correlated high-pass signal, which is then added to the narrowband secondary channel in the band combiner 934 to obtain the wideband secondary channel, and finally the wideband primary channel and the wideband secondary channel are output.

[0046] FIG. 11 shows a further embodiment in which a decoded base channel obtained by a base channel decoder 700 of a particular sampling rate for the encoded base channel is input to a resampler 710 to obtain a resampled base channel, which is then used in a multi-channel processor operating on the resampled channel.

[0047] 12 shows a preferred implementation of reference stereo encoding. In block 1200, an inter-channel phase difference IPD is calculated for a first channel, such as L, and a second channel, such as R. This IPD value is then typically quantized and output for each band in each time frame as encoder output data 1206. Furthermore, the IPD value is calculated for each time frame. Each band of TIFF0007804634000002.tif32 Prediction parameters for TIFF0007804634000003.tif43 TIFF0007804634000004.tif47 and each time frame Each band of TIFF0007804634000005.tif32 Gain parameters for TIFF0007804634000006.tif43 Used to calculate parametric data for stereo signals, such as TIFF0007804634000007.tif46.

[0048] Additionally, both the first and second channels are also used in the mid / side processor 1203 to calculate the mid and side signals for each band.

[0049] Depending on the implementation, the intermediate signal Only TIFF0007804634000008.tif44 can be sent to encoder 1204, and the side signal is not sent to encoder 1204, so that output data 1206 includes only the encoded base channel, the parametric data generated by block 1202, and the IPD information generated by block 1200.

[0050] It should be noted that although the preferred embodiment will now be described with respect to the reference encoder, any other stereo encoder such as those mentioned above may be used as well.

[0051] Reference Stereo Encoder We will consider a DFT-based stereo encoder as a reference. As usual, we use the time-frequency vectors L for the left and right channels. t and R t is generated by simultaneously applying an analysis window and then a Discrete Fourier Transform (DFT). The DFT bins are then divided into subbands (L t,k ) k ∈ I b and (R t,k ) k ∈ I b where I b denotes a set of subband indices.

[0052] IPD calculation and downmixing. For downmixing, the inter-channel phase difference (IPD) for each band is calculated as follows:

[0053] (1) Calculated as TIFF0007804634000009.tif652, where TIFF0007804634000010.tif45 is This means the complex conjugate of TIFF0007804634000011.tif33. Intermediate signal for each band for TIFF0007804634000012.tif512

[0054] (2) TIFF0007804634000013.tif1053 and side signal

[0055] (3) used to generate TIFF0007804634000014.tif1052, Here, β is, for example,

[0056] (4) The absolute phase rotation parameters are given by TIFF0007804634000015.tif990.

[0057] Parameter calculation. In addition to the per-band IPD, two further stereo parameters are extracted: By TIFF0007804634000016.tif59 TIFF0007804634000017.tif57, i.e., the optimal coefficient for predicting the residual energy

[0058] (5) The number for which TIFF0007804634000018.tif541 is the smallest TIFF0007804634000019.tif47, and intermediate signal When applied to TIFF0007804634000020.tif56, the TIFF0007804634000021.tif45 and Relative gain factor to equalize the energy of TIFF0007804634000022.tif56 TIFF0007804634000023.tif46, i.e.

[0059] (6) TIFF0007804634000024.tif1535Optimal prediction coefficients are calculated based on the subband energy

[0060] (7) TIFF0007804634000025.tif738 and TIFF0007804634000026.tif737 and TIFF0007804634000027.tif55 and Absolute value of the dot product of TIFF0007804634000028.tif55

[0061] (8) From TIFF0007804634000029.tif648,

[0062] (9) It can be calculated as TIFF0007804634000030.tif947.

[0063] from now, TIFF0007804634000031.tif47 is in [-1, 1]. The residual gain is calculated from the energy and dot product.

[0064] (10) TIFF0007804634000032.tif46= TIFF0007804634000033.tif1168 can be calculated in the same way, which is

[0065] (11) means TIFF0007804634000034.tif1041.

[0066] A preferred embodiment on the decoder side is shown in Figure 13. In block 700, which corresponds to the base channel decoder of Figure 7a, the encoded base channel TIFF0007804634000035.tif44 is decoded.

[0067] Next, in block 940a, the primary upmix channels are calculated, such as L. Furthermore, in block 940b, the primary upmix channels are calculated, such as L. The secondary upmix channel is calculated as TIFF0007804634000036.tif43.

[0068] Both blocks 940a and 940b are connected to the supplemental signal generator 800 and receive the parametric data generated by block 1200 of FIG. 12 or 1202 of FIG.

[0069] Preferably, the parametric data is provided in bands having a second spectral resolution, and blocks 940a, 940b operate at a higher spectral resolution granularity to generate spectral lines at a first spectral resolution that is higher than the second spectral resolution.

[0070] The outputs of blocks 940a, 940b are, for example, inputs to frequency-to-time transformers 961, 962. These transformers may be DFTs or any other transforms, and typically also include subsequent synthesis windowing and further overlap-add operations.

[0071] In addition, the fill signal generator receives an energy normalization coefficient, preferably a compressed energy normalization coefficient, which is used to generate fill signal spectral lines of correct levels / weighting for blocks 940a and 940b.

[0072] Next, preferred implementations of blocks 940a and 940b are shown. Both blocks include a calculation 941a of phase rotation coefficients and a calculation of first weights for the spectral lines of the decoded base channel as indicated by 942a and 942b. Furthermore, both blocks include calculations 943a and 943b for calculating second weights for the spectral lines of the supplemental signal.

[0073] Additionally, the fill signal generator 800 receives the energy normalization coefficients generated by block 945. This block 945 receives the fill signals per band and the base channel signals per band, and then calculates the same energy normalization coefficient to be used for all lines within a band.

[0074] Finally, this data is transferred to a processor 946 for calculating the spectral lines of the first and second upmix channels. To this end, the processor 946 receives the data from blocks 941a, 941b, 942a, 942b, 943a, and 943b, as well as the spectral lines of the decoded base channel and the spectral lines of the filler signal. The output of block 946 is therefore the corresponding spectral lines of the first and second upmix channels.

[0075] Next, a preferred embodiment of the decoder is presented.

[0076] Reference Decoder We describe a reference DFT-based decoder corresponding to the encoders mentioned above. The time-frequency transforms from both encoders are applied to the decoded downmix, resulting in a time-frequency vector TIFF0007804634000037.tif69 is obtained. Dequantized values TIFF0007804634000038.tif612, TIFF0007804634000039.tif57, and Using TIFF0007804634000040.tif56, the left and right channels are About TIFF0007804634000041.tif511

[0077] (12) TIFF0007804634000042.tif965 and

[0078] (13) TIFF0007804634000043.tif975, where TIFF0007804634000044.tif57 is the missing residual from the encoder It is an alternative to TIFF0007804634000045.tif47, TIFF0007804634000046.tif412 is the energy normalization coefficient

[0079] (14) TIFF0007804634000047.tif1132 and the relative residual prediction gain Convert TIFF0007804634000048.tif46 to absolute gain. A quick selection about TIFF0007804634000049.tif57 is:

[0080] (15) TIFF0007804634000050.tif730, where TIFF0007804634000051.tif510 means frame delay per band, but this has certain drawbacks: · TIFF0007804634000052.tif44 and TIFF0007804634000053.tif55 may have very different spectral and temporal shapes, Even in the case of harmonic spectral and temporal envelopes, the use of (15) in (12) and (13) leads to frequency-dependent ILDs and IPDs that vary only slowly in the low-to-mid frequency range, which causes problems for e.g. tonal items, For speech signals, the delay should be chosen small to stay below the echo threshold, but this causes strong tones due to comb filtering.

[0081] Therefore, it is better to use the time-frequency bins of the artificial signal, which will be described later.

[0082] The phase rotation coefficient β is also

[0083] (16) Calculated as TIFF0007804634000054.tif989.

[0084] Synthetic signal generation A second signal is added to the time-domain input signal to replace the missing residual part in the stereo upmix. The second signal generated from tif44 The output is TIFF0007804634000056.tif57. The design constraint of this filter is to have a short and dense impulse response. This is achieved by applying several stages of a basic all-pass filter, which is obtained by nesting two Schroeder all-pass filters within a third Schroeder filter, i.e.

[0085] (17) TIFF0007804634000057.tif650, where

[0086] (18) TIFF0007804634000058.tif1032TIFF0007804634000059.tif1017 and

[0087] (19) The file is TIFF0007804634000060.tif930.

[0088] These basic all-pass filters

[0089] (20) TIFF0007804634000061.tif913 has been proposed by Schroeder in the context of artificial reverberation generation, and is applied with both large gains and large delays. Since it is undesirable in this context to have a reverberant output signal, the gains and delays are chosen to be fairly small. As in the case of reverberation, a dense random-like impulse response is applied with pairs of disjoint delays for all all-pass filters. This is best achieved by selecting TIFF0007804634000062.tif54.

[0090] The filter operates at a fixed sampling rate, regardless of the bandwidth or sampling rate of the signal provided by the core coder. When used with an EVS coder, this is necessary because the bandwidth may be changed by the bandwidth detector during operation, and a fixed sampling rate ensures a consistent output. The preferred sampling rate for the all-pass filter is 32 kHz, the native ultra-wideband sampling rate, because the absence of residual components above 16 kHz is usually no longer audible. When used with an EVS coder, the signal is constructed directly from the core, which includes several resampling routines, as shown in Figure 1.

[0091] Filters that have been shown to work well at 32kHz sampling rates are

[0092] (twenty one) TIFF0007804634000063.tif536, where TIFF0007804634000064.tif54 is a basic all-pass filter with the gain and delay shown in Table 1. The impulse response of this filter is shown in Figure 6. For complexity reasons, such filters can also be applied at lower sampling rates and / or with fewer basic all-pass filter units.

[0093] The all-pass filter unit also provides the ability to overwrite parts of the input signal with zeros, controlled by the encoder, which can be used, for example, to remove attacks from the filter input.

[0094] coefficient Compressing TIFF0007804634000065.tif412 Energy adjustment gain that compresses values ​​towards 1 for a smoother output Applying a compressor to TIFF0007804634000066.tif412 has proven beneficial, as it also compensates a bit for the fact that some of the ambience is typically lost after coding the downmix at a lower bitrate.

[0095] Such a compressor is

[0096] (twenty two) This can be constructed by taking tif0007804634000067.tif554, where

[0097] (twenty three) TIFF0007804634000068.tif739 and the function TIFF0007804634000069.tif33

[0098] (twenty four) Fill TIFF0007804634000070.tif525.

[0099] next, Around TIFF0007804634000071.tif32 The value in TIFF0007804634000072.tif33 specifies how strongly this region will be compressed, with a value of 0 corresponding to no compression and a value of 1 corresponding to full compression. Additionally, the compression scheme can be TIFF0007804634000073.tif33 is even, i.e. In the case of TIFF0007804634000074.tif526, it is symmetric. An example is

[0100] (twenty five) TIFF0007804634000075.tif948, which is

[0101] (26) Resulting in TIFF0007804634000076.tif560.

[0102] In this case, (22) is

[0103] (27) This can be simplified to TIFF0007804634000077.tif5105, saving extra function evaluations. Using Bandwidth Extension for ACELP Frames in Combination with Time-Domain Stereo Upmix When used with the EVS codec, a low-delay audio codec in the communications context, it is desirable to perform a stereo upmix of the bandwidth extension in the time domain to accommodate the delay caused by the time-domain bandwidth extension (TBE). The stereo bandwidth upmix aims to restore the correct panning in the bandwidth extension range, but does not add a replacement for the missing residual. Therefore, it is desirable to add a replacement in the frequency-domain stereo processing, as shown in Figure 2.

[0104] Input signal for decoder TIFF0007804634000078.tif44, About the filtered input signal TIFF0007804634000079.tif57, About the time-frequency bins in TIFF0007804634000080.tif44 TIFF0007804634000081.tif69, and About the time-frequency bins in TIFF0007804634000082.tif57 The notation TIFF0007804634000083.tif57 is used.

[0105] next, We are facing the problem that TIFF0007804634000084.tif69 is unknown in the bandwidth extension range, so the index Energy normalization factor when part of TIFF0007804634000085.tif59 is in the bandwidth extension range

[0106] (28) TIFF0007804634000086.tif1541 cannot be calculated directly. This problem is solved as follows: TIFF0007804634000087.tif46 and Let TIFF0007804634000088.tif56 represent the high and low band indices of the frequency bins, respectively. Then, Rating for TIFF0007804634000089.tif726 TIFF0007804634000090.tif511 is obtained by calculating the energy of the windowed highband signal in the time domain, where TIFF0007804634000091.tif59 and TIFF0007804634000092.tif510 is the bandwidth This is the index of TIFF0007804634000093.tif43 If we represent the low and high band indexes in TIFF0007804634000094.tif44,

[0107] (29) TIFF0007804634000095.tif724= The file is TIFF0007804634000096.tif762.

[0108] where the addend in the second sum on the right is unknown, TIFF0007804634000097.tif57 is filtered by an all-pass filter Because it is obtained from TIFF0007804634000098.tif44, TIFF0007804634000099.tif57 and The energy of TIFF0007804634000100.tif59 can be assumed to be similarly distributed, and therefore,

[0109] (30) It is thought to be TIFF0007804634000101.tif1282.

[0110] Therefore, the second sum on the right side of (29) is

[0111] (31) It can be evaluated as TIFF0007804634000102.tif1122TIFF0007804634000103.tif829.

[0112] Use in coders coding primary and secondary channels Artificial signals are also useful in stereo coders that code primary and secondary channels. In this case, the primary channel serves as the input to an all-pass filter unit. The filtered output, possibly after applying a shaping filter, can then be used to replace the residual portion of the stereo processing. In the simplest configuration, the primary and secondary channels can be transforms of the input channels, such as mid / side or KL transforms, and the secondary channel can be limited to a smaller bandwidth. The missing portion of the secondary channel can then be replaced by the filtered primary channel after applying a high-pass filter.

[0113] For use in decoders that can switch between stereo modes A particularly interesting case of artificial signals is when the decoder is equipped with different stereo processing methods, as shown in Figure 3. These methods may be applied simultaneously (e.g., separated by bandwidth) or exclusively (e.g., frequency domain vs. time domain processing) and connected in a switching decision. Using the same artificial signal for all stereo processing methods smooths discontinuities in both the switching and simultaneous cases.

[0114] Benefits and Advantages of the Preferred Embodiments This novel method has a number of benefits and advantages compared to state-of-the-art methods applied in, for example, xHE-AAC.

[0115] Time-domain processing allows for much higher temporal resolution than subband processing, which, when applied to parametric stereo, makes it possible to design filters with dense and fast-decaying impulse responses, making the spectral envelope of the input signal less likely to distort over time or color the output signal, and therefore sounding more natural.

[0116] To be more suitable for speech, the optimum peak region of the filter's impulse response should be located between 20 and 40 milliseconds.

[0117] The filter unit is capable of resampling input signals with different sampling rates. This allows the filter to operate at a fixed sampling rate, which is beneficial because it ensures similar outputs at different sampling rates or smooths discontinuities when switching between signals with different sampling rates. For complexity reasons, the internal sampling rate should be chosen so that the filtered signal covers only the perceptually important frequency range.

[0118] Because the signal is generated at the decoder input and is not connected to a filter bank, it can be used in different stereo processing units, which helps to smooth out discontinuities when switching between different units or when operating different units on different parts of the signal.

[0119] It also reduces complexity as no reinitialization is required when switching between units.

[0120] The gain compression scheme helps to compensate for the loss of atmosphere caused by core coding.

[0121] The method related to the bandwidth extension of ACELP frames mitigates the lack of missing residual components in the panning-based time-domain bandwidth extension upmix, which increases the stability when switching between high-band processing in the DFT domain and the time domain.

[0122] The input can be replaced by zeros on a very fine time scale, which is useful in handling attacks.

[0123] Further details regarding FIG. 1a or 1b, FIG. 2a or 2b, and FIG. 3 will now be described.

[0124] 1a or 1b shows a base channel decoder 700 comprising a first decoding branch having a low-band decoder 721 and a bandwidth extension decoder 720 for generating a first portion of a decoded base channel. Furthermore, the base channel decoder 700 comprises a second decoding branch 722 having a full-band decoder for generating a second portion of the decoded base channel.

[0125] The switching between both elements is performed by a controller 713, shown as a switch controlled by a control parameter contained in the encoded multi-channel signal, to feed part of the encoded base channel either to a first decoding branch comprising blocks 720, 721 or to a second decoding branch 722. The low-band decoder 721 is realized, for example, as an algebraic code excited linear predictive coder ACELP, while the second full-band decoder is realized as a transform coded excitation (TCX) / high quality (HQ) core decoder.

[0126] The decoded downmix from block 722 or the decoded core signal from block 721, as well as the bandwidth extension signal from block 720, are taken and sent to the procedure of Figure 2a or 2b. Further, a subsequent decorrelation filter comprises resamplers 810, 811, 812 and, if necessary, delay compensation elements 813, 814. An adder combines the time-domain bandwidth extension signal from block 720 with the core signal from block 721 and sends it to a switch 815 controlled by the encoded multi-channel data in the form of a switch controller for switching between the first coding branch or the second coding branch depending on which signal is available.

[0127] Further, a switch decision 817 is configured, which may for example be implemented as a transient detector. However, the transient detector does not necessarily have to be an actual detector that detects transients by signal analysis, but the transient detector may also be configured to determine certain control parameters of the encoded multi-channel signal that are indicative of transients in the side information or base channel.

[0128] A switch decision 817 sets the switch to either provide the signal output from switch 815 to all-pass filter unit 802 or provide a zero input which actually stops the addition of filler signals in the multi-channel processor for a particular very specific selectable time region because the EVS all-pass signal generator (APSG) shown at 1000 in Figure 1a or 1b operates entirely in the time domain. Thus, the zero input can be selected in a sample-by-sample manner without needing to reference a window length which reduces spectral resolution as required for spectral domain processing.

[0129] The apparatus shown in Figure 1a differs from that shown in Figure 1b in that the resampler and delay stages have been omitted in Figure 1b, i.e. elements 810, 811, 812, 813, 814 are not necessary in the apparatus of Figure 1b. Thus, in the embodiment of Figure 1b, the all-pass filter unit operates at 16 kHz rather than 32 kHz as in Figure 1a.

[0130] 2a or 2b show the integration of an all-pass signal generator 1000 into DFT stereo processing including a time-domain bandwidth extension upmix. Block 1000 outputs the bandwidth extension signal generated by block 720 to a high-band upmixer 960 (TBE upmix—(time-domain) bandwidth extension upmix) to generate a high-band left signal and a high-band right signal from the mono bandwidth extension signal generated by block 720. Further, a resampler 821 is provided and connected before the DFT for the filler signal indicated by 804. In addition, a DFT 922 for the decoded base channel, which is either a (full-band) decoded downmix or a (low-band) decoded core signal, is provided.

[0131] Depending on the implementation, if a decoded downmix signal from the fullband decoder 722 is available, block 960 is stopped and the stereo processing block 904 already outputs a fullband upmix signal, such as fullband left and right channels.

[0132] However, when the decoded core signal is input to DFT block 922, block 960 is activated and the left and right channel signals are added by adders 994a and 994b. However, the addition of the filler signal is still performed in the spectral domain indicated by block 904, for example, according to the procedure as described in the preferred embodiment based on Equations 28-31. Therefore, in such a situation, the signal output by DFT block 902 corresponding to the low-band intermediate signal does not have high-band data. However, the signal output by block 804, i.e., the filler signal, has low-band data and high-band data.

[0133] In the stereo processing block, the low-band data output by block 904 is generated by the decoded base channel and the filler signal, while the high-band data output by block 904 consists only of the filler signal and does not have the high-band information from the decoded base channel due to the band-limited nature of the decoded base channel. The high-band information from the decoded base channel is generated by the bandwidth extension block 720 and upmixed by block 960 to the left high-band channel and the right high-band channel, and then added by summers 994a and 994b.

[0134] The apparatus shown in Figure 2a differs from the apparatus shown in Figure 2b in that the resampler has been omitted in Figure 2b, ie element 821 is not required in the apparatus of Figure 2b.

[0135] 3 shows a preferred embodiment of a system having multiple stereo processing units 904a, 904b, 904c as described above with respect to switching between stereo modes. Each stereo processing block receives side information and also a particular primary signal, but also receives the exact same filler signal, regardless of whether a particular time portion of the input signal is processed using stereo processing algorithm 904a, stereo processing algorithm 904b, or another stereo processing algorithm 904c.

[0136] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of a corresponding method, with blocks or apparatus corresponding to method steps or features of method steps. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps can be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or electronic circuitry. In some embodiments, one or more of the most important method steps can be performed by such an apparatus.

[0137] The encoded audio signals of the present invention can be stored on a digital storage medium or transmitted over a transmission medium, such as a wireless or wired transmission medium, such as the Internet.

[0138] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. Implementation can be performed using non-transitory or digital storage media such as floppy disks, DVDs, Blu-rays, CDs, ROMs, PROMs, EPROMs, EEPROMs, or flash memories that store electronically readable control signals and that cooperate (or can cooperate) with a computer system that can be programmed to perform the respective methods. Thus, the digital storage media can be computer-readable.

[0139] Some embodiments according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.

[0140] Generally, embodiments of the present invention can be realized as a computer program product having program code operable to perform one of the above methods when the computer program product is run on a computer. The program code can, for example, be stored on a machine-readable carrier.

[0141] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0142] In other words, therefore, one embodiment of the inventive methods is a computer program having a program code for performing one of the methods described herein when the computer program runs on a computer.

[0143] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium or computer readable medium) having recorded thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium or recorded medium is typically tangible and / or non-transitory.

[0144] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, The data stream or the sequence of signals can be adapted to be transmitted via a data communication connection, for example the Internet.

[0145] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0146] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0147] Further embodiments according to the invention include an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, comprise a file server for transmitting the computer program to the receiver.

[0148] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0149] The apparatus described herein can be implemented using a hardware apparatus, a computer, or a combination of a hardware apparatus and a computer.

[0150] The devices described herein, or any components of the devices described herein, may be implemented at least in part in hardware and / or software.

[0151] The methods described herein can be performed using a hardware apparatus, using a computer, or using a combination of a hardware apparatus and a computer.

[0152] The methods described herein, or any components of the methods described herein, may be performed at least in part by hardware and / or software.

[0153] The above-described embodiments are merely illustrative of the principles of the present invention. It is to be understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. Therefore, the present invention is limited only by the scope of the appended claims, and not by the specific details presented by way of illustration and description of the embodiments herein.

[0154] In the foregoing description, it will be appreciated that various features are grouped together in the embodiments for the purpose of streamlining the disclosure. This method of disclosure should not be interpreted as reflecting an intention that the claimed embodiments require additional features other than those expressly recited in each claim. Rather, as the appended claims reflect, inventive subject matter may include less than all features of a single disclosed embodiment. Accordingly, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separate embodiment. While each claim may stand on its own as a separate embodiment, a dependent claim may relate to specific combinations with one or more other claims in the claim, and it should be noted that other embodiments may include combinations of a dependent claim with the subject matter of each other dependent claim, or combinations of each feature with other dependent or independent claims. Such combinations are suggested herein unless it is expressly stated that a specific combination is not intended. Furthermore, incorporating features of one claim into another independent claim is contemplated even if that claim is not directly dependent on that independent claim.

[0155] Furthermore, it should be noted that the methods disclosed herein or in the claims may be performed by an apparatus having means for performing each of the respective steps of those methods.

[0156] Furthermore, in some embodiments, a single step may include or be divided into multiple sub-steps, and such sub-steps may be included in and be part of the disclosure of the single step, unless expressly excluded.

Claims

1. 1. An apparatus for decoding an encoded multi-channel signal, comprising: a base channel decoder (700) for decoding the encoded base channel to obtain a decoded base channel; a decorrelation filter (800) for filtering at least a portion of the decoded base channel to obtain a supplemental signal; a multi-channel processor (900) for performing multi-channel processing using the spectral representation of the decoded base channel and the spectral representation of the supplemental signal; It is equipped with the decorrelation filter (800) is a wideband filter, and the multi-channel processor (900) is configured to apply narrowband processing to the spectral representation of the decoded base channel and the spectral representation of the supplemental signal; The multi-channel processor (900) configured to calculate an energy normalization factor using spectral bands of the decoded base channel and corresponding spectral bands of the supplemental signal; and determining (946) a first upmix channel and a second upmix channel using different weighting combinations of the spectral bands of the decoded base channel and the corresponding spectral bands of the supplemental signal, the different weighting combinations depending on the energy normalization coefficient. Device.

2. the spectral representation of the decoded base channel has a first spectral granularity indicating bandwidths associated with individual spectral lines of the spectral representation of the decoded base channel, and the spectral representation of the filler signal has a second spectral granularity indicating bandwidths associated with individual spectral lines of the spectral representation of the filler signal; a filter characteristic of the decorrelation filter (800) having an area of ​​constant size, the decorrelation filter (800) being configured such that the area of ​​constant size is greater in frequency than the first spectral granularity of the spectral representation of the decoded base channel and the second spectral granularity of the spectral representation of the supplemental signal; 10. The apparatus of claim 1.

3. The decorrelation filter (800) a filter stage (802) for filtering the decoded base channel to obtain a wideband or time-domain supplemental signal; a spectral transformer (804) for transforming the wideband or time-domain supplemental signal into the spectral representation of the supplemental signal; 3. The apparatus according to claim 1 or 2, comprising:

4. a base channel spectral converter (902) for converting the decoded base channel into the spectral representation of the decoded base channel. An apparatus according to any one of claims 1 to 3.

5. the decorrelation filter (800) comprises an all-pass time-domain filter (802) or at least one Schroeder all-pass filter (802); An apparatus according to any one of claims 1 to 4.

6. the multi-channel processor (900) is configured to compress the energy normalization coefficients to obtain compressed energy normalization coefficients and to calculate the different weighting combinations using the compressed energy normalization coefficients. An apparatus according to any one of claims 1 to 5.

7. The energy normalization factor is Calculating the logarithm of the energy normalization factor (921); subjecting the logarithm to a non-linear function (922); applying an exponential function to the result of the non-linear function (923); The apparatus of claim 6, wherein the compressed data is compressed using

8. The nonlinear function is where c is a function, The function c is is defined based on where t is a real number and τ is the integration variable.

8. The apparatus of claim 7.

9. the multi-channel processor (900, 924, 925) is configured to compress (921) the energy normalization coefficients to obtain compressed energy normalization coefficients, and to use the compressed energy normalization coefficients to calculate the different weighting combinations using a non-linear function; The nonlinear function is is defined based on where α is a predetermined boundary value and t is a value between −α and +α.

10. The device of claim 1 or 7.

10. the multi-channel processor (900) is configured to calculate (904) a low-band first upmix channel and a low-band second upmix channel; The apparatus further comprises a time-domain bandwidth extender (960) for extending the low-band first upmix channel and the low-band second upmix channel or the low-band base channel; the energy normalization factor is calculated using the energy of the spectral bands of the decoded base channel and the corresponding spectral bands of the supplemental signal; The energy normalization factor is calculated (961) using an energy estimate derived from the energy of the windowed highband signal. An apparatus according to any one of claims 1 to 9.

11. the time domain bandwidth extender (960) is configured to use a high-band signal without a windowing operation used in the calculation of the energy normalization factor.

11. The apparatus of claim 10.

12. the base channel decoder (700, 705) is configured to provide a decoded primary base channel and a decoded secondary base channel; the decorrelation filter (800) is configured to filter the decoded primary base channel to obtain the supplemental signal; the multi-channel processor (900) is configured to perform multi-channel processing by synthesizing one or more residual portions of the multi-channel processing using the supplemental signals; a shaping filter (930) is applied to the supplemental signal; An apparatus according to any one of claims 1 to 11.

13. the primary and secondary base channels are the result of a transformation of the original input channels, the transformation being, for example, a mid-side transform or a Karhunen-Loeve (KL) transform, and the decoded secondary base channel is limited to a smaller bandwidth; the multi-channel processor (900) is configured to high-pass filter (930) the supplemental signal and use the high-pass filtered supplemental signal as a secondary channel for bandwidths not included in the bandwidth-limited decoded secondary base channel.

13. The apparatus of claim 12.

14. the multi-channel processor (900) is configured to perform different multi-channel processing methods (904a, 904b, 904c); the multi-channel processor (900) performs the different multi-channel processing methods simultaneously, separated by bandwidth, or exclusively connected to a switching decision, a first multi-channel processing method among the different multi-channel processing methods including frequency domain processing, and a second multi-channel processing method among the different multi-channel processing methods including time domain processing; the multi-channel processor (900) is configured to use the same supplementary signal in the first and second multi-channel processing methods (904a, 904b, 904c). An apparatus according to any one of claims 1 to 13.

15. The decorrelation filter (800) is a time domain filter (802) having an optimum peak region of a time domain filter impulse response between 20 ms and 40 ms. An apparatus according to any one of claims 1 to 14.

16. the decorrelation filter (800) is configured to resample (811, 812) the decoded base channel in time portions to a predetermined target sampling rate or to an input-dependent target sampling rate; the decorrelation filter (800) is configured to filter the resampled decoded base channel using a decorrelation filter (802) stage; the multi-channel processor (900) is configured to convert (710) the decoded base channel for the further time portion to the predetermined target sampling rate or the input-dependent target sampling rate, so that the multi-channel processor (900) operates using a spectral representation of the decoded base channel and the supplemental signal that is based on the predetermined target sampling rate or the input-dependent target sampling rate, despite different sampling rates of the decoded base channel for the time portion and the further time portion; or The device is configured to perform resampling before, during, or after the transformation (804, 702) into the frequency domain.

16. Apparatus according to any one of claims 1 to 15.

17. The base channel decoder (700) a first decoding branch comprising a low-band decoder (721) and a bandwidth extension decoder (720) for generating a first portion of the decoded channel; a second decoding branch (722) having a full-band decoder and generating a second portion of the decoded base channel; a controller (713) for providing a portion of the encoded base channel to either the first decoding branch or the second decoding branch in response to a control signal; Equipped with An apparatus according to any one of claims 1 to 16.

18. The decorrelation filter (800) a first resampler (810, 811) for resampling the first portion to a predetermined sampling rate; a second resampler (812) for resampling the second portion to the predetermined sampling rate; an all-pass filter unit (802) for all-pass filtering an all-pass filter input signal to obtain said supplemental signal; a controller (815) for supplying the resampled first portion or the resampled second portion to the all-pass filter unit (802); 18. The apparatus of any one of claims 1 to 17, comprising:

19. the controller (815) is configured to provide either the resampled first portion, the resampled second portion, or zero data (816) to the all-pass filter unit in response to a control signal; 20. The apparatus of claim 18.

20. The decorrelation filter (800) a time-to-spectral converter (804) for converting said supplemental signal into a spectral representation comprising spectral lines of a first spectral resolution; Equipped with the multi-channel processor (900) comprises a time-to-spectral converter (902) for converting the decoded base channels into a spectral representation using spectral lines of the first spectral resolution; the multi-channel processor (904) is configured to generate spectral lines having the first spectral resolution for the first upmix channel or the second upmix channel using, for a particular spectral line, the spectral lines of the supplemental signal, the spectral lines of the decoded base channel, and one or more parameters; the one or more parameters relate to a second spectral resolution that is lower than the first spectral resolution; the one or more parameters are used to generate a set of spectral lines including the particular spectral line and at least one frequency-adjacent spectral line; 20. Apparatus according to any one of claims 1 to 19.

21. The multi-channel processor may generate spectral lines for the first upmix channel or the second upmix channel as: Phase rotation coefficients (941a, 941b) depending on one or more transmitted parameters; the spectral lines of the decoded base channel; first weights (942a, 942b) of the spectral lines of the decoded base channel according to transmitted parameters; the spectral lines of said replenishment signal; second weights (943a, 943b) of the spectral lines of the supplemental signal according to the transmitted parameters; and the energy normalization factor 12. The apparatus of claim 1, configured to generate a signal using:

22. for the calculation of the second upmix channel, the signs of the second weights are different from the signs of the second weights used for the calculation of the first upmix channel, or for the calculation of the second upmix channel, the phase rotation coefficients are different from the phase rotation coefficients used for the calculation of the first upmix channel; for the calculation of the second upmix channel, the first weights are different from the first weights used for the calculation of the first upmix channel; 22. The apparatus of claim 21.

23. the different weighting combinations further depend on prediction coefficients and gain factors, and the apparatus is configured to decode encoded parameters comprised in the encoded multi-channel signal to obtain the prediction coefficients and the gain factors.

10. The apparatus of claim 1.

24. 1. A method for decoding an encoded multi-channel signal, comprising: decoding (700) the encoded base channel to obtain a decoded base channel; decorrelation filtering (800) at least a portion of the decoded base channel to obtain a supplemental signal; performing multi-channel processing (900) using a spectral representation of the decoded base channel and a spectral representation of the supplemental signal; It contains The decorrelation filtering (800) is a wideband filtering and the multi-channel processing (900) is calculating an energy normalization factor using the spectral bands of the decoded base channel and the corresponding spectral bands of the supplemental signal; applying narrowband processing to the spectral representation of the decoded base channel and the spectral representation of the supplemental signal; the step of performing multi-channel processing (900) includes determining (946) a first up-mix channel and a second up-mix channel using different weighting combinations of the spectral bands of the decoded base channel and the corresponding spectral bands of the filler signal, the different weighting combinations depending on the energy normalization factor; method.

25. 25. A computer program for performing the method of claim 24 when run on a computer or processor.

Citation Information

Patent Citations

  • Signal synthesis method

    JP2005523624A

  • JPP7161233B