Multichannel audio decoder, multichannel audio encoder, method, and computer program using residual signal-based adjustment of the contribution of uncorrelated signal

JP2026139655APending Publication Date: 2026-09-01FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2026078075
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2013-10-18
Filing Date
2026-05-07
Publication Date
2026-09-01

AI Technical Summary

Benefits of technology

【0013】 好ましい実施形態において、マルチチャンネルオーディオデコーダは、(また)無相関化信号に従って、重み付け結合における無相関化信号の寄与を記述する重みを決定するように構成される。残差信号と無相関化信号の両方に従って、重み付け結合における無相関化信号の寄与を記述する重みを決定することによって、符号化表現に基づいて(特に、ダウンミックス信号と無相関化信号と残差信号とに基づいて)、少なくとも2つの出力オーディオ信号の良好な品質の復元を達成することができるように、信号特性に対して重みを適切に調整することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026139655000001_ABST
    Figure 2026139655000001_ABST
Patent Text Reader

Abstract

It provides efficient encoding and decoding of multi-channel audio signals. [Solution] A multichannel audio decoder 200 that provides at least two output audio signals based on an encoded representation performs a weighted combination 220 of a downmix signal, an uncorrelated signal, and a residual signal to obtain one of the output audio signals and determines weights 232 that describe the contribution of the uncorrelated signal in the weighted combination according to the residual signal. A multichannel audio encoder that provides an encoded representation of a multichannel audio signal obtains a downmix signal based on the multichannel audio signal, provides parameters that describe the inter-channel dependencies, provides a residual signal, and changes the amount of residual signal included in the encoded representation according to the multichannel audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Embodiments of the present invention relate to a multi-channel audio decoder that provides at least two output audio signals based on an encoded representation.

[0002] Another embodiment of the present invention relates to a multichannel audio encoder that provides an encoded representation of a multichannel audio signal.

[0003] Another embodiment of the present invention relates to a method for providing at least two output audio signals based on an encoded representation.

[0004] Another embodiment of the present invention relates to a method for providing an encoded representation of a multichannel audio signal.

[0005] Another embodiment of the present invention relates to a computer program that performs one of the above methods.

[0006] In general, some embodiments of the present invention relate to combined residual coding and parametric coding. [Background technology]

[0007] In recent years, the demand for storage and transmission of audio content has steadily increased. Furthermore, the demand for quality in the storage and transmission of audio content has also steadily increased. Therefore, concepts for encoding and decoding audio content have been strengthened. For example, the so-called "Advanced Audio Coding" (AAC), described in Non-Patent Document 1, has been developed.

[0008] Furthermore, several spatial extensions have been developed, for example, such as the so-called "MPEG Surround" concept described in Non-Patent Document 2. Further, Non-Patent Document 3 describes additional improvements to the encoding and decoding of spatial information of audio signals relating to so-called spatial audio object coding. Furthermore, Non-Patent Document 4 defines a flexible (switchable) audio encoding / decoding concept that describes the so-called "united speech and audio coding (USAC)" concept, encodes both general audio signals and speech signals with good coding efficiency, and provides the possibility to handle multi-channel audio signals. [Prior Art Document] [Non-Patent Document]

[0009] [Non-Patent Document 1] International Standard ISO / IEC 13818-7:2003 [Non-Patent Document 2] International Standard ISO / IEC 23003-1:2007 [Non-Patent Document 3] International Standard ISO / IEC 23003-2:2010 [Non-Patent Document 4] International Standard ISO / IEC 23003-3:2012 [Summary of the Invention] [Problem to be Solved by the Invention]

[0010] However, there is a demand for providing even more advanced concepts for efficient encoding and decoding of multi-channel audio signals. [Means for Solving the Problem]

[0011] Embodiments of the present invention construct a multichannel audio decoder that provides at least two output audio signals based on an encoded representation. The multichannel audio decoder is configured to obtain one of the output audio signals by performing a weighted combination of a downmix signal, a decorrelated signal, and a residual signal. The multichannel audio decoder is configured to determine weights that describe the contribution of the decorrelated signal in the weighted combination according to the residual signal.

[0012] This embodiment of the present invention is based on the discovery that an output audio signal can be obtained in a highly efficient manner based on the encoded representation when the weights describing the contribution of the uncorrelated signal to the weighted coupling of the downmix signal, the uncorrelated signal, and the residual signal are adjusted according to the residual signal. Thus, by adjusting the weights describing the contribution of the uncorrelated signal in the weighted coupling according to the residual signal, it is possible to mix (or fade) between parametric coding (or primarily parametric coding) and residual coding (or mostly residual coding) without transmitting additional control information. Furthermore, it is generally preferable to give the uncorrelated signal a (relatively) high weight when the residual signal is (relatively) weak (or insufficient for restoring the desired energy), and a (relatively) small weight when the residual signal is (relatively) strong (or sufficient for restoring the desired energy), so the residual signal included in the encoded representation has been found to be a good indicator for the weights describing the contribution of the uncorrelated signal in the weighted coupling. Therefore, the above concept allows for a stepwise transition between parametric coding (e.g., desired energy and / or correlation characteristics are signaled by parameters and restored by adding an uncorrelated signal) and residual coding (the residual signal is used to restore the output audio signal—and possibly the waveform of the output audio signal—based on the downmix signal). Thus, it is possible to adapt the restoration technique and the quality of the restoration to the decoded signal without the overhead of additional signaling.

[0013] In a preferred embodiment, the multichannel audio decoder is configured to determine weights describing the contribution of the uncorrelated signal in the weighted coupling, according to the uncorrelated signal. By determining weights describing the contribution of the uncorrelated signal in the weighted coupling, according to both the residual signal and the uncorrelated signal, the weights can be appropriately adjusted to the signal characteristics so that a good quality restoration of at least two output audio signals can be achieved based on the encoded representation (in particular, based on the downmix signal, the uncorrelated signal, and the residual signal).

[0014] In a preferred embodiment, the multichannel audio decoder is configured to obtain upmix parameters based on the encoded representation and to determine weights that describe the contribution of the uncorrelated signals in the weighted combination according to the upmix parameters. By considering the upmix parameters, it is possible to reconstruct desired characteristics of the output audio signals (such as desired correlation between output audio signals and / or desired energy characteristics of the output audio signals) to take desired values.

[0015] In a preferred embodiment, the multichannel audio decoder is configured to determine weights describing the contribution of the uncorrelated signal in the weighted coupling such that the weight of the uncorrelated signal decreases with increasing energy of one or more residual signals. This mechanism allows for adjusting the accuracy of the reconstruction of at least two output audio signals according to the energy of the residual signals. When the energy of the residual signals is relatively high, the weight of the contribution of the uncorrelated signal is relatively small so that the uncorrelated signal does not have a detrimental effect on the high quality of reproduction that would result from using the residual signals. In contrast, when the energy of the residual signals is relatively low or zero, a high weight is given to the uncorrelated signal so that the uncorrelated signal can efficiently bring the characteristics of the output audio signals to the desired value.

[0016] In a preferred embodiment, the multichannel audio decoder is configured to determine weights describing the contribution of the uncorrelated signal in the weighted combination such that when the energy of the residual signal is zero, the maximum weight determined by the uncorrelated signal upmix parameter is associated with the uncorrelated signal, and when the energy of the residual signal weighted using the residual signal weight coefficient is greater than or equal to the energy of the uncorrelated signal weighted by the uncorrelated signal upmix parameter, the zero weight is associated with the uncorrelated signal. This embodiment is based on the finding that the desired energy to be added to the downmix signal is determined by the energy of the uncorrelated signal weighted by the uncorrelated signal upmix parameter. Thus, it can be concluded that the uncorrelated signal no longer needs to be added when the energy of the residual signal weighted by the residual signal weight coefficient is greater than or equal to the energy of the uncorrelated signal weighted by the uncorrelated signal upmix parameter. In other words, when it is determined that the residual signal has sufficient energy (e.g., sufficient to reach a sufficient total energy), the uncorrelated signal is no longer used to provide at least two output audio signals.

[0017] In a preferred embodiment, the multichannel audio decoder is configured to determine a factor according to the weighted energy values ​​of the uncorrelated signal and the weighted energy values ​​of the residual signal, and to calculate the weighted energy values ​​of the uncorrelated signal weighted according to one or more uncorrelated signal upmix parameters in order to obtain weights describing the contribution of the uncorrelated signal to (at least) one audio output signal based on that factor, and to calculate the weighted energy values ​​of the residual signal weighted using one or more residual signal upmix parameters (which may be equal to the residual signal weighting factors described above). This procedure has been found to be well suited to the efficient calculation of weights describing the contribution of the uncorrelated signal to one or more output audio signals.

[0018] In a preferred embodiment, the multichannel audio decoder is configured to multiply the coefficients by the uncorrelated signal upmix parameter to obtain weights describing the contribution of the uncorrelated signal to (at least) one output audio signal. Using such a procedure, it is possible to consider both one or more parameters describing the desired signal characteristics of at least two output audio signals (described by the uncorrelated signal upmix parameter) and the relationship between the energy of the uncorrelated signal and the energy of the residual signal in order to determine the weights describing the contribution of the uncorrelated signal in the weighted combination. Thus, it is possible to mix (or fade) between parametric coding (or primarily parametric coding) and residual coding (or primarily residual coding) while considering the desired characteristics of the output audio signals (which are reflected by the uncorrelated signal upmix parameter).

[0019] In a preferred embodiment, the multichannel audio decoder is configured to obtain a weighted energy value of the uncorrelated signal by calculating the energy of the uncorrelated signal weighted using the uncorrelated signal upmix parameter across multiple upmix channels and time slots. Thus, it is possible to avoid large changes in the weighted energy value of the uncorrelated signal. Consequently, stable tuning of the multichannel audio decoder is achieved.

[0020] Similarly, a multi-channel audio decoder is configured to obtain a weighted energy value for the residual signal by calculating the energy of the residual signal weighted using the residual signal upmix parameters across multiple upmix channels and time slots. Thus, stable tuning of the multi-channel audio decoder is achieved, as large changes in the weighted energy value of the residual signal are avoided. However, the averaging period can be selected to be sufficiently short to allow for dynamic adjustment of the weighting.

[0021] In a preferred embodiment, the multichannel audio decoder is configured to calculate coefficients according to the difference between the weighted energy values ​​of the uncorrelated signal and the weighted energy values ​​of the residual signal. The calculation comparing the weighted energy values ​​of the uncorrelated signal and the weighted energy values ​​of the residual signal allows for the supplementation of the residual signal (or a weighted version of the residual signal) with the uncorrelated signal (or a weighted version of the residual signal), and the weights are adjusted to describe the contribution of the uncorrelated signal to the need to provide at least two audio channel signals.

[0022] In a preferred embodiment, the multichannel audio decoder is configured to calculate a coefficient according to the ratio of the difference between the weighted energy value of the uncorrelated signal and the weighted energy value of the residual signal to the weighted energy value of the uncorrelated signal. Calculation of the coefficient according to this ratio has been shown to yield long-term, characteristically good results. Furthermore, it should be noted that this ratio describes which portion of the total energy value of the uncorrelated signal (weighted using the uncorrelated signal upmix parameter) is necessary in the presence of the residual signal to achieve a good auditory impression (or, equivalently, to have substantially the same signal energy in the output audio signal compared to the case without the residual signal).

[0023] In a preferred embodiment, the multichannel audio decoder is configured to determine weights describing the contribution of the uncorrelated signal to two or more output audio signals. In this case, the multichannel audio decoder is configured to determine the contribution of the uncorrelated signal to a first output audio signal based on the weighting energy value of the uncorrelated signal and the uncorrelated signal upmix parameter of the first channel. Furthermore, the multichannel audio decoder is configured to determine the contribution of the uncorrelated signal to a second output audio channel based on the weighting energy value of the uncorrelated signal and the uncorrelated signal upmix parameter of the second channel. Thus, the two output audio signals can be delivered with reasonable effort and good audio quality, and the difference between the two output audio signals is taken into account by using the uncorrelated signal upmix parameter of the first channel and the uncorrelated signal upmix parameter of the second channel.

[0024] In a preferred embodiment, the multichannel audio decoder is configured to disable the contribution of the uncorrelated signal to weighted coupling when the residual energy exceeds the energy of the uncorrelatedizer (i.e., the energy of the uncorrelated signal or its weighted version). Thus, when the residual signal has sufficient energy and the residual energy exceeds the energy of the uncorrelatedizer, it is possible to switch to pure residual coding without using the uncorrelated signal.

[0025] In a preferred embodiment, the audio decoder is configured to determine, band by band, weights describing the contribution of the uncorrelated signal in the weighted coupling, according to the band-by-band determination of the weighted energy values ​​of the residual signal. Thus, it is possible to flexibly determine, without additional signaling overhead, which frequency bands should be based on (or primarily on) parametric coding for the improvement of at least two output audio signals, and which frequency bands should be based on (or primarily on) residual coding for the improvement of at least two output audio signals. Thus, it is possible to flexibly determine which frequency bands should be used (or at least primarily on) waveform reconstruction (or at least partial waveform reconstruction) using residual coding while keeping the weight of the uncorrelated signal relatively small. Thus, good audio quality can be obtained by selectively applying parametric coding (which is primarily based on the supply of the uncorrelated signal) and residual coding (which is primarily based on the supply of the residual signal).

[0026] In a preferred embodiment, the audio decoder is configured to determine weights for each frame of the output audio signal that describe the contribution of the uncorrelated signal in the weighted coupling. This allows for fine timing resolution and flexible switching between parametric coding (or primarily parametric coding) and residual coding (or primarily residual coding) between subsequent frames. Thus, audio decoding can be adjusted with good temporal resolution to the characteristics of the audio signal.

[0027] Another embodiment of the present invention constructs a multichannel audio decoder that provides at least two output audio signals based on an encoded representation. The multichannel audio decoder is configured to obtain (at least) one output audio signal based on an encoded representation of a downmix signal and an encoded representation of a plurality of encoded spatial parameters and residual signals. The multichannel audio decoder is configured to mix between parametric encoding and residual encoding according to the residual signal. Thus, a highly flexible audio decoding concept is achieved in which the best decoding mode (parametric encoding / decoding vs (versus) residual encoding / decoding) can be selected without additional signaling overhead. Furthermore, the considerations described above also apply.

[0028] Embodiments of the present invention construct a multichannel audio encoder that provides an encoded representation of a multichannel audio signal. The multichannel audio encoder is configured to acquire a downmix signal based on the multichannel audio signal. Furthermore, the multichannel audio encoder is configured to provide parameters that describe the inter-channel dependencies of the multichannel audio signal and to provide residual signals. Furthermore, the multichannel audio encoder is configured to vary the amount of residual signal included in the encoded representation according to the multichannel audio signal. By varying the amount of residual signal included in the encoded representation, the encoding process can be flexibly adjusted to the characteristics of the signal. For example, it is possible to include a relatively large amount of residual signal in the encoded representation for parts of the waveform of the decoded audio signal where it is desirable to preserve at least partially (e.g., the temporal part and / or the frequency part). Thus, the possibility of varying the amount of residual signal included in the encoded representation enables a more accurate residual-signal-based reconstruction of the multichannel audio signal. Furthermore, it should be noted that the above-described multichannel audio decoder does not even require additional signaling for a mixture between (primarily) parametric encoding and (primarily) residual encoding, thus constructing a very efficient concept in combination with the above-described multichannel audio decoder. Therefore, the multi-channel encoder described here makes it possible to take advantage of the benefits that become available by using the multi-channel audio encoder described above.

[0029] In a preferred embodiment, the multichannel audio encoder is configured to vary the bandwidth of the residual signal according to the multichannel audio signal. Thus, it is possible to adjust the residual signal so that it helps to restore the psychoacoustically most important frequency band or frequency range.

[0030] In a preferred embodiment, the multichannel audio encoder is configured to select the frequency bands to which the residual signal should be included in the encoded representation, according to the multichannel audio signal. Thus, the multichannel audio encoder can determine which frequency bands are necessary or most beneficial to include the residual signal (which typically results in at least partial waveform reconstruction). For example, psychoacoustically significant frequency bands can be considered. Additionally, the presence of transient events can also be considered, as the residual signal typically helps improve the rendering of transient phenomena in the audio decoder. Furthermore, the available bitrate can also be taken into account to determine how much residual signal should be included in the encoded representation.

[0031] In a preferred embodiment, the multichannel audio encoder is configured to selectively include residual signals in the encoded representation for frequency bands where the multichannel audio signal is tonal, while excluding the inclusion of residual signals in the encoded representation for frequency bands where the multichannel audio signal is not tonal. This embodiment is based on the consideration that the audio quality obtainable on the audio decoder side can be improved, especially when the tonal frequency bands are reproduced with particularly high quality, preferably when waveform reconstruction is used at least partially. Therefore, selectively including residual signals in the encoded representation for frequency bands where the multichannel audio signal is tonal is beneficial because it results in a good compromise between bitrate and audio quality.

[0032] In a preferred embodiment, the multichannel audio encoder is configured to selectively include residual signals in the encoded representation for the time portion and / or frequency bands resulting from the cancellation of signal components of the multichannel audio signal when forming the downmix signal. It has been found that when there is cancellation of components in the multichannel audio signal, it is difficult or even impossible to properly reconstruct multiple audio signals based on the downmix signal, since the canceled signal components cannot be recovered even through uncorrelated or predictive means when forming the downmix signal. In such cases, the use of residual signals is an effective way to avoid significant degradation of the reconstructed multichannel audio signal. Thus, this concept helps to improve audio quality while avoiding signaling effort (for example, when combined with the audio decoder described above).

[0033] In a preferred embodiment, the multichannel audio encoder is configured to detect the cancellation of signal components of the multichannel audio signal in the downmix signal, and the multichannel audio decoder is configured to activate the provision of a residual signal in response to the detection result. Thus, there is an effective way to avoid poor audio quality.

[0034] In a preferred embodiment, the multichannel audio encoder is configured to calculate a residual signal using a linear combination of at least two channel signals of the multichannel audio signal and its dependence on the upmix coefficient used on the multichannel decoder side. Thus, the residual signal is calculated in an efficient manner and is well-suited for the reconstruction of the multichannel audio signal on the multichannel audio decoder side.

[0035] In one embodiment, the multichannel audio encoder is configured to encode upmix coefficients using parameters that describe the inter-channel dependencies of the multichannel audio signal, or to derive upmix coefficients from parameters that describe the inter-channel dependencies of the multichannel audio signal. Thus, the provision of residual signals can be efficiently performed based on parameters that are also used for parametric coding.

[0036] In a preferred embodiment, the multichannel audio encoder is configured to use a psychoacoustic model to determine the amount of residual signal included in the encoded representation as a time variable. Therefore, a relatively high amount of residual signal can be included for portions of the multichannel audio signal with relatively high psychoacoustic relevance (time portions, frequency portions, or time-frequency portions), while a (relatively) smaller amount of residual signal can be included for time portions, frequency portions, or time-frequency portions of the multichannel audio signal with relatively low psychoacoustic relevance. Thus, a good trade-off between bitrate and audio quality can be achieved.

[0037] In a preferred embodiment, the multichannel audio encoder is configured to determine, as a time variable, the amount of residual signal included in the encoded representation according to the currently available bitrate. Thus, the audio quality can be adapted to the available bitrate, making it possible to obtain the best possible audio quality for the currently available bitrate.

[0038] Embodiments of the present invention construct a method for providing at least two output audio signals based on an encoded representation. The method comprises the step of obtaining one of the output audio signals by performing a weighted combination of a downmix signal, an uncorrelated signal, and a residual signal. The weights describing the contribution of the uncorrelated signal in the weighted combination are determined according to the residual signal. This method is based on the same considerations as the audio decoder described above.

[0039] Another embodiment of the present invention constructs a method for providing at least two output audio signals based on an encoded representation. The method comprises the step of obtaining (at least) one output audio signal based on an encoded representation of a downmix signal and an encoded representation of a residual signal with a plurality of encoded spatial parameters. Mixing (or fading) is performed between parametric encoding and residual encoding according to the residual signal. This method is also based on the same considerations as the audio decoder described above.

[0040] Another embodiment of the present invention constructs a method for providing an encoded representation of a multichannel audio signal. The method comprises the steps of: obtaining a downmix signal based on a multichannel audio signal; providing parameters describing the inter-channel dependencies of the multichannel audio signal; and providing a residual signal. The amount of residual signal included in the encoded representation is varied according to the multichannel audio signal. This method is based on the same considerations as the audio encoder described above.

[0041] A further embodiment of the present invention involves constructing a computer program that performs the method described in the present specification. [Brief explanation of the drawing]

[0042] Embodiments relating to the present invention will be described subsequently with reference to the following drawings. [Figure 1] Figure 1 shows a schematic block diagram of a multi-channel audio encoder according to one embodiment of the present invention. [Figure 2] Figure 2 shows a schematic block diagram of a multi-channel audio decoder according to one embodiment of the present invention. [Figure 3] Figure 3 shows a schematic block diagram of a multichannel audio decoder according to another embodiment of the present invention. [Figure 4]Figure 4 shows a flowchart of a method for providing an encoded representation of a multichannel audio signal according to one embodiment of the present invention. [Figure 5] Figure 5 shows a flowchart of a method for providing at least two output audio signals based on an encoded representation according to one embodiment of the present invention. [Figure 6] Figure 6 shows a flowchart of a method for providing at least two output audio signals based on an encoded representation according to another embodiment of the present invention. [Figure 7] Figure 7 shows a flowchart of a decoder according to one embodiment of the present invention. [Figure 8] Figure 8 shows a schematic representation of the hybrid residual decoder. [Modes for carrying out the invention]

[0043] 1. Multichannel audio encoder shown in Figure 1

[0044] Figure 1 shows a schematic block diagram of a multi-channel audio encoder 100 that provides an encoded representation of a multi-channel signal.

[0045] The multichannel audio encoder 100 is configured to receive a multichannel audio signal 110 and provide an encoded representation 112 of the multichannel audio signal 110 based on it. The multichannel audio encoder 100 also includes a processor (or processing device) 120 configured to receive a multichannel audio signal and obtain a downmix signal 122 based on the multichannel audio signal 110. The processor 120 is further configured to provide parameters 124 that describe the inter-channel dependencies of the multichannel audio signal 110. Furthermore, the processor 120 is configured to provide a residual signal 126. The multichannel audio encoder also includes residual signal processing 130 configured to vary the amount of residual signal contained in the encoded representation 112 according to the multichannel audio signal 110.

[0046] However, it should be noted that a multi-channel audio decoder does not necessarily need to have a separate processor 120 and a separate residual signal processing unit 130. Rather, it is sufficient if the multi-channel audio encoder is configured in some way to perform the functions of the processor 120 and the residual signal processing unit 130.

[0047] Regarding the functionality of the multichannel audio encoder 100, it should be noted that the channel signals of the multichannel audio signal 110 are typically encoded using multichannel coding, and the encoded representation 112 typically comprises a downmix signal 122 (in encoded form), parameters 124 describing the dependencies between the channels (or channel signals) of the multichannel audio signal 110, and a residual signal 126. The downmix signal 122 can be based, for example, on a combination (e.g., a linear combination) of the channel signals of the multichannel audio signal. However, the downmix signal 122 can be provided based on multiple channel signals of the multichannel audio signal. However, two or more downmix signals can relate to a larger number of channel signals (usually more than the number of downmix signals) of the multichannel audio signal 110. Parameters 124 can describe the dependencies (e.g., correlation, covariance, level relationships, etc.) between the channels (or channel signals) of the multichannel audio signal 110. Therefore, parameter 124 serves the purpose of the audio decoder deriving a restored version of the channel signals of the multichannel audio signal 110 based on the downmix signal 122. To this end, parameter 124 describes the desired characteristics (e.g., individual or relative characteristics) of the channel signals of the multichannel audio signal so that an audio encoder using parametric decoding can restore the channel signals based on one or more downmix signals 122.

[0048] In addition, the multi-channel audio decoder 100 provides a residual signal 126 that typically represents signal components that cannot be restored by the audio decoder (e.g., an audio decoder following specific processing rules) based on the downmix signal 122 and parameters 124, as predicted or estimated by the multi-channel audio encoder. Therefore, the residual signal 126 can typically be considered an improved signal that enables waveform restoration, or at least partial waveform restoration, on the audio decoder side.

[0049] However, the multichannel audio encoder 100 is configured to vary the amount of residual signal included in the encoded representation 112 according to the multichannel audio signal 110. In other words, the multichannel audio encoder can determine, for example, the intensity (or energy) of the residual signal 126 included in the encoded representation 112. In addition, or alternatively, the multichannel audio encoder 100 can determine for which frequency bands and / or how many frequency bands the residual signal is included in the encoded representation 112. By varying the "amount" of residual signal 126 included in the encoded representation 112 according to the multichannel audio signal (and / or the available bitrate), the multichannel audio encoder 100 can flexibly determine with what accuracy the channel signals of the multichannel audio signal 110 can be reconstructed on the audio decoder side based on the encoded representation 112. Thus, the accuracy with which the channel signals of the multichannel audio signal 110 can be reconstructed can be adapted to the psychoacoustic relationships of different signal parts of the channel signals of the multichannel audio signal 110 (e.g., time part, frequency part and / or time / frequency part). Therefore, signal portions with high psychoacoustic relevance (e.g., tonal signal portions or signal portions with transient events) can be encoded with particularly high resolution by including a "large" amount of residual signal 126 in the encoded representation. For example, it can be achieved that residual signals with relatively high energy are included in the encoded representation 112 for signal portions with high psychoacoustic relevance. Furthermore, if the downmix signal 122 contains "low quality" signals, for example, when the channel signals of a multi-channel audio signal 112 are coupled to the downmix signal 122, and there is substantial cancellation of signal components, it can be achieved that high-energy residual signals are included in the encoded representation 112.In other words, the multi-channel audio decoder 100 can selectively embed a "larger amount" of residual signal (e.g., a residual signal with relatively high energy) into the encoded representation of the signal portion of the multi-channel audio signal 110, where providing a relatively large amount of residual signal results in a significant improvement in the restored channel signal (restored by the audio decoder).

[0050] Therefore, the change in the amount of residual signal included in the encoded representation according to the multichannel audio signal 110 makes it possible to adapt the encoded representation 112 of the multichannel audio signal 110 (for example, the residual signal 126 included in the encoded representation in encoded form) so that a good trade-off can be achieved between bitrate efficiency and the audio quality of the restored multichannel audio signal (restored on the audio decoder side).

[0051] It should be noted that the multi-channel audio encoder 100 can be improved as an option in many different ways. For example, the multi-channel audio encoder can be configured to vary the bandwidth of the residual signal 126 (included in the encoded representation) according to the multi-channel audio signal 110. Thus, the amount of residual signal included in the encoded representation 112 can be adapted to the perceptually most important frequency band.

[0052] Optionally, the multichannel audio decoder can be configured to select the frequency bands in which the residual signal 126 is contained in the encoded representation 112 according to the multichannel audio signal 110. Thus, the encoded representation 120 (more precisely, the amount of residual signal contained in the encoded representation 112) can be adapted to the multichannel audio signal, for example, to the perceptually most important frequency bands of the multichannel audio signal 110.

[0053] Optionally, the multichannel audio encoder can be configured to include residual signals 126 in the encoded representation for frequency bands where the multichannel audio signal is tonal. In addition, the multichannel audio encoder can be configured not to include residual signals 126 in the encoded representation 112 for frequency bands where the multichannel audio signal is not tonal (unless any other specific conditions are met that would result in the inclusion of residual signals in the encoded representation for a particular frequency band). Thus, residual signals can be selectively included in the encoded representation for perceptually important tonal frequency bands.

[0054] Optionally, the multichannel audio encoder 100 can be configured to selectively include residual signals in the encoded representation for time portions and / or frequency bands where the formation of a downmix signal results in the cancellation of signal components of the multichannel audio signal. For example, the multichannel audio encoder can be configured to detect the cancellation of signal components of the multichannel audio signal 110 in the downmix signal 122 and, in response to the result of the detection, activate the provision of residual signals 126 (e.g., inclusion of residual signals 126 in the encoded representation 112). Thus, if the downmix of the channel signals of the multichannel audio signal 110 to the downmix signal 122 (or any other normal linear combination) results in the cancellation of signal components of the multichannel audio signal 112 (which may be caused, for example, by signal components of different channel signals that are 180 degrees phase-shifted), the encoded representation 112 includes residual signals 126 that help overcome the detrimental effects of this cancellation when the audio decoder restores the multichannel audio signal 110. For example, the residual signal 126 can be selectively included in the encoded representation 112 for frequency bands where such cancellations occur.

[0055] Optionally, the multichannel audio encoder can be configured to calculate the residual signal using a linear combination of at least two channel signals of the multichannel audio signal, according to the upmix coefficient used on the multichannel audio decoder side. Such calculation of the residual signal is efficient and allows for easy reconstruction of the channel signals on the audio decoder side.

[0056] Optionally, the multichannel audio encoder can be configured to encode upmix coefficients using parameter 124 that describes the inter-channel dependencies of the multichannel audio signal, or to derive upmix coefficients from parameters that describe the inter-channel dependencies of the multichannel audio signal. Thus, parameter 124 (which can be, for example, an intra-channel level difference parameter, an intra-channel correlation parameter, etc.) can be used for both parametric coding (encoding or decoding) and residual signal-assisted coding (encoding or decoding). Therefore, the use of residual signal 126 does not result in additional signaling overhead. Rather, parameter 124 used for parametric coding (encoding / decoding) in either case is reused for residual coding (encoding / decoding). Therefore, high coding efficiency can be achieved.

[0057] Optionally, a multi-channel audio decoder can be configured to use a psychoacoustic model to determine the amount of residual signal included in the encoded representation as a time variable. Thus, the encoding accuracy can be adapted to the psychoacoustic characteristics of the signal, which usually results in good bitrate efficiency.

[0058] However, it should be noted that a multi-channel audio encoder can be optionally supplemented by any of the features or functions described in this specification (both the specification and the claims). Furthermore, a multi-channel audio encoder can be adapted in parallel with the audio decoder described in this specification in order to work with the audio decoder.

[0059] 2. Multichannel audio decoder shown in Figure 2

[0060] Figure 2 shows a schematic block diagram of a multi-channel audio decoder 200 according to one embodiment of the present invention.

[0061] The multichannel audio decoder 200 is configured to receive an encoded representation 210 and provide at least two output audio signals 212, 214 based thereon. The multichannel audio decoder 200 may include, for example, a weighted coupler 220 configured to perform a weighted coupling of a downmix signal 222, an uncorrelated signal 224, and a residual signal 226 to obtain (at least) one output signal, for example, a first output audio signal 212. Here, it should be noted that the downmix signal 222, the uncorrelated signal 224, and the residual signal 226 can be derived, for example, from the encoded representation 210, and the encoded representation 210 may be accompanied by an encoded representation of the downmix signal 220 and an encoded representation of the residual signal 226. Furthermore, the uncorrelated signal 224 can be derived, for example, from the downmix signal 222, or can be derived using additional information contained in the encoded representation 210. However, the uncorrelated signal can also be provided without any dedicated information from the encoded representation 210.

[0062] The multichannel audio decoder 200 is also configured to determine weights describing the contribution of the uncorrelated signal 224 in the weighted coupling, according to the residual signal 226. For example, the multichannel audio decoder 200 may include a weight determination unit 230 configured to determine weights 232 describing the contribution of the uncorrelated signal 224 in the weighted coupling (e.g., the contribution of the uncorrelated signal 224 to the first output audio signal 212) based on the residual signal 226.

[0063] Regarding the functionality of the multi-channel audio decoder 200, it should be noted that the contribution of the uncorrelated signal 224 to the weighted coupling, and consequently to the first output audio signal 212, is adjusted in a flexible manner (e.g., in a time-variable and frequency-dependent manner) according to the residual signal 226 without additional signaling overhead. Therefore, the amount of uncorrelated signal 224 included in the first output audio signal 212 is adapted according to the amount of residual signal 226 included in the first output audio signal 212 so that good quality of the first output audio signal 212 is achieved. Thus, it is possible to obtain appropriate weighting of the uncorrelated signal 224 without additional signaling overhead under any circumstances. Consequently, good quality of the decoded output audio signal 212 can be achieved at a moderate bitrate using the multi-channel audio decoder 200. The accuracy of the reconstruction can be flexibly adjusted by the audio encoder, which can determine the amount of residual signal 226 contained in the encoded representation 210 (e.g., how large the energy of the residual signal 226 contained in the encoded representation 210 is, or to what frequency band the residual signal 226 contained in the encoded representation 210 relates to), and the multi-channel audio decoder 200 can react accordingly and adjust the weighting of the uncorrelated signal 224 to fit the amount of residual signal 226 contained in the encoded representation 210. Consequently, when there is a large amount of residual signal 226 contained in the encoded representation 210 (e.g., for a particular frequency band or for a particular time portion), the weighted coupling 220 can primarily (or exclusively) consider the residual signal 226, while little (or no) weight is given to the uncorrelated signal 224. In contrast, when the encoded representation 210 contains only a smaller amount of residual signal 226, the weighted concatenation 220 can consider the downmix signal 222, as well as primarily (or exclusively) the uncorrelated signal 224, but the residual signal 226 is given only a relatively small amount of weight (or no weight at all).Therefore, the multi-channel audio decoder 200 can flexibly work with a suitable multi-channel audio encoder and adjust the weighted coupling 220 to achieve the best audio quality under any circumstances (regardless of whether a smaller or larger amount of residual signal 226 is included in the encoded representation 210).

[0064] It should be noted that the second output audio signal 214 can be generated in a similar manner. However, if there are different quality requirements for the second output audio signal, for example, it is not always necessary to apply the same mechanism to the second output audio signal 214.

[0065] In an optional improvement, the multi-channel audio decoder can be configured to determine weights 232 that describe the contribution of the uncorrelated signal 224 to the weighted coupling, according to the uncorrelated signal 224. In other words, weights 232 can be dependent on both the residual signal 226 and the uncorrelated signal 224. Thus, weights 232 can even be better fitted to the current decoded audio signal without additional signaling overhead.

[0066] As an improvement to other options, the multichannel audio decoder can be configured to obtain upmix parameters based on the encoded representation 212 and determine weights 232 that describe the contribution of the uncorrelated signals in the weighted combination according to the upmix parameters. Thus, the weights 232 can be additionally dependent on the upmix parameters to achieve a better fit of the weights 232.

[0067] As an improvement to other options, the multichannel audio decoder can be configured to determine weights describing the contribution of the uncorrelated signal in the weighted coupling such that the weight of the uncorrelated signal decreases with increasing energy of the residual signal. Thus, mixing or fading can be performed between decoding based primarily on the uncorrelated signal 224 (in addition to the downmix signal 222) and decoding based primarily on the residual signal 226 (in addition to the downmix signal 222).

[0068] As an improvement to other options, the multi-channel audio decoder 200 can be configured to determine the weights 232 such that when the energy of the residual signal 226 is zero, the maximum weight determined by the uncorrelated signal upmix parameter (which may be included in or derived from the coded representation 210) relates to the uncorrelated signal 224, and when the energy of the residual signal 226 weighted by the residual signal weight coefficient (or residual signal upmix parameter) is greater than or equal to the energy of the uncorrelated signal 224 weighted by the uncorrelated signal upmix parameter, the zero weight relates to the uncorrelated signal. Thus, it is possible to completely blend (or fade) between decoding based on the uncorrelated signal 224 and decoding based on the residual signal 226. If the residual signal 226 is judged to be sufficiently strong (for example, when the energy of the weighted residual signal is equal to or greater than the energy of the weighted uncorrelated signal 224), the weighted coupling can be made to rely entirely on the residual signal 226 to improve the downmix signal 222, without taking the uncorrelated signal 224 into consideration. In this case, the use of the residual signal 226 usually enables good waveform reconstruction, whereas the consideration of the uncorrelated signal 224 usually hinders particularly good waveform reconstruction, so that particularly good (at least partial) waveform reconstruction can be performed on the multichannel audio decoder 200 side.

[0069] In other optional improvements, the multichannel audio decoder 200 can be configured to calculate the weighted energy value of the uncorrelated signal weighted according to one or more uncorrelated signal upmix parameters, and to calculate the weighted energy value of the residual signal weighted using one or more residual signal upmix parameters. In this case, the multichannel audio decoder can be configured to determine coefficients according to the weighted energy values ​​of the uncorrelated signal and the weighted energy values ​​of the residual signal, and to obtain weights based on these coefficients that describe the contribution of the uncorrelated signal 224 to one output audio signal (e.g., a first output audio signal 212). Thus, the weight determination 230 can provide particularly well-fitted weight values ​​232.

[0070] In an optional improvement, the multichannel audio decoder 200 (or its weighting determiner 230) may be configured to multiply its coefficients by an uncorrelated signal upmix parameter (which may be included in or derived from the encoded representation 210) in order to obtain a weight (or weighting value) 232 that describes the contribution of the uncorrelated signal 224 to one output audio signal (e.g., a first output audio signal 212).

[0071] In an optional improvement, the multichannel audio decoder (or its weighting determiner 230) may be configured to calculate the energy of the uncorrelated signal 224 weighted with uncorrelated signal upmix parameters (which may be included in or derived from the encoded representation 210) across multiple upmix channels and time slots in order to obtain the weighted energy value of the uncorrelated signal 224.

[0072] As a further option for improvement, the multi-channel audio decoder 200 can be configured to calculate the energy of a weighted residual signal 226 using residual signal upmix parameters (which may be included in or derived from the encoded representation 210) across multiple upmix channels and time slots in order to obtain weighted energy values ​​of the residual signal.

[0073] As an improvement to other options, the multi-channel audio decoder 200 (or its weighting determiner 230) can be configured to calculate the coefficients described above according to the difference between the weighted energy values ​​of the uncorrelated signal and the weighted energy values ​​of the residual signal. Such calculations have been found to be an efficient solution for determining the weighted values ​​232.

[0074] As an optional improvement, the multichannel audio decoder can be configured to calculate a coefficient according to the ratio of the difference between the weighted energy value of the uncorrelated signal 224 and the weighted energy value of the residual signal 226 to the weighted energy value of the uncorrelated signal 224. Such calculations for the coefficient have been shown to yield good results for a blend of primarily uncorrelated signal-based improvements to the downmix signal 222 and primarily residual signal-based improvements to the downmix signal 222.

[0075] As an optional improvement, the multi-channel audio decoder 200 can be configured to determine weights describing the contribution of the uncorrelated signal to two or more output audio signals, such as a first output audio signal 212 and a second output audio signal 214. In this case, the multi-channel audio decoder can be configured to determine the contribution of the uncorrelated signal 224 to the first output audio signal 212 based on the weighted energy value of the uncorrelated signal 224 and the uncorrelated signal upmix parameter of the first channel. Furthermore, the multi-channel audio decoder can be configured to determine the contribution of the uncorrelated signal 224 to the second output audio signal 214 based on the weighted energy value of the uncorrelated signal 224 and the uncorrelated signal upmix parameter of the second channel. In other words, different uncorrelated signal upmix parameters can be used to provide the first output audio signal 212 and the second output audio signal 214. However, the same weighted energy value of the uncorrelated signal can be used to determine the contribution of the uncorrelated signal to the first output audio signal 212 and the contribution of the uncorrelated signal to the second output audio signal 214. Therefore, effective adjustments can be made that can be considered by different uncorrelated signal upmix parameters, regardless of the different characteristics of the two output audio signals 212 and 214.

[0076] As an optional improvement, the multi-channel audio decoder 200 can be configured to disable the contribution of the uncorrelated signal 224 to the weighted coupling if the residual energy (e.g., the energy of the residual signal 226 or the energy of the weighted version of the residual signal 226) exceeds the uncorrelated energy (e.g., the energy of the uncorrelated signal 224 or the energy of the weighted version of the uncorrelated signal 224).

[0077] As a further option for improvement, the audio decoder can be configured to determine, band by band, weights 232 describing the contribution of the uncorrelated signal 224 in the weighted coupling, according to the band by band determination of the weighted energy values ​​of the residual signal. Thus, fine-tuning of the multi-channel audio decoder 200 to the signal being decoded can be performed.

[0078] In other optional improvements, the audio decoder can be configured to determine weights for each frame of the output audio signals 212,214 that describe the contribution of the uncorrelated signals in the weighted combination. Thus, good temporal resolution can be achieved.

[0079] In further refinement of the options, the determination of the weight value 232 can be carried out by several formulas provided below.

[0080] Furthermore, it should be noted that the multi-channel audio decoder 200 can be supplemented with any of the features or functions described in this specification in other embodiments as well.

[0081] 3. Multichannel audio decoder shown in Figure 3

[0082] Figure 3 shows a schematic block diagram of a multichannel audio decoder 300 according to one embodiment of the present invention. The multichannel audio decoder 300 is configured to receive an encoded representation 310 and provide two or more output audio signals 312, 314 based thereon. The encoded representation 310 may include, for example, an encoded representation of a downmix signal, an encoded representation of one or more spatial parameters, and an encoded representation of a residual signal. The multichannel audio decoder 300 is configured to obtain (at least) one output audio signal, for example, a first output audio signal 312 and / or a second output audio signal 314, based on the encoded representation of the downmix signal, a plurality of encoded spatial parameters, and an encoded representation of a residual signal.

[0083] In particular, the multichannel audio decoder 300 is configured to mix between parametric coding and residual coding according to a residual signal (which is included in the coded form in the coded representation 310). In other words, the multichannel audio decoder 300 can mix between a decoding mode in which the output audio signals 312,314 are provided using a downmix signal and spatial parameters (e.g., a desired inter-channel level difference or a desired inter-channel correlation of the output audio signals 312,314) that describe a desired relationship between the output audio signals 312,314, and a decoding mode in which the output audio signals 312,314 are restored based on the downmix signal using a residual signal. Therefore, the intensity (e.g., energy) of the residual signal included in the encoded representation 310 can determine whether the decoding is based solely (or exclusively) on spatial parameters (in addition to the downmix signal) or solely (or exclusively) on the residual signal (in addition to the downmix signal) in order to derive the output audio signals 312,314 from the downmix signal, or whether an intermediate state is taken in which both spatial parameters and the residual signal affect the improvement of the downmix signal.

[0084] Furthermore, the multi-channel audio decoder 300 enables decoding that is well-suited to the current audio content without high signaling overhead by mixing parametric coding (typically with relatively high weights given to the uncorrelated signal when providing output audio signals 312,314) and residual coding according to the residual signal (typically with relatively low weights given to the uncorrelated signal).

[0085] Furthermore, it should be noted that the multi-channel audio decoder 300 is based on similar considerations to the multi-channel audio decoder 200, and the optional improvements described above for the multi-channel audio decoder 200 can also be applied to the multi-channel audio decoder 300.

[0086] 4. A method for providing an encoded representation of a multichannel audio signal as shown in Figure 4.

[0087] Figure 4 shows a flowchart of method 400, which provides an encoded representation of a multichannel audio signal.

[0088] Method 400 comprises step 410 of obtaining a downmix signal based on a multichannel audio signal. Method 400 comprises step 420 of providing parameters that describe the inter-channel dependencies of the multichannel audio signal. For example, inter-channel level difference parameters and / or inter-channel correlation parameters (or covariance parameters) that describe the inter-channel dependencies of the multichannel audio signal can be provided. Method 400 also comprises step 430 of providing a residual signal. Furthermore, Method 440 of varying the amount of residual signal included in the encoded representation according to the multichannel audio signal.

[0089] It should be noted that Method 400 is based on the same considerations as the audio encoder 100 shown in Figure 1. Furthermore, Method 400 can be supplemented with any of the features and functions described in this specification with respect to the apparatus of the invention.

[0090] 5. A method for providing at least two output audio signals based on the encoded representation shown in Figure 5.

[0091] Figure 5 shows a flowchart of method 500 for providing at least two output audio signals based on an encoded representation. Method 500 comprises a step 510 for determining weights that describe the contribution of the uncorrelated signal in the weighted combination according to the residual signal. Method 500 also comprises a step 520 for obtaining one of the output audio signals by performing a weighted combination of the downmix signal, the uncorrelated signal and the residual signal.

[0092] It should be noted that Method 500 can be supplemented by any of the features and functions described in the present specification with respect to the apparatus of the present invention.

[0093] 6. A method for providing at least two output audio signals based on the encoded representation shown in Figure 6.

[0094] Figure 6 shows a flowchart of method 600 for providing at least two output audio signals based on encoded representations. Method 600 comprises step 610 of obtaining one of the output audio signals based on an encoded representation of a downmix signal and an encoded representation of a residual signal with a plurality of encoded spatial parameters. Step 610 of obtaining one of the output audio signals comprises step 620 of performing a blend between parametric encoding and residual encoding according to the residual signal.

[0095] It should be noted that Method 600 can be supplemented by any of the features and functions described in the present specification with respect to the apparatus of the present invention.

[0096] 7. Further Embodiments

[0097] Some general considerations and some further embodiments are described below.

[0098] 7.1 General Considerations

[0099] Embodiments of the present invention are based on the idea that, instead of using a fixed residual bandwidth, the decoder (e.g., a multi-channel audio decoder) detects the amount of transmitted residual signal by measuring the energy band by band for each frame (or, more generally, for at least several frequency ranges and / or several time segments). Depending on the transmitted spatial parameters, the uncorrelated output is added where the residual energy is "lost" to obtain the output energy and the required (or desired) amount of uncorrelatedness. This allows for variable residual bandwidth, as well as bandpass-style residual signals. For example, it is possible to use residual coding only for tonal bands. A residual signal for a simplified downmix is ​​defined in this specification to allow the use of a simplified downmix for parametric coding, as is the case for waveform-preserving coding (also known as residual coding).

[0100] 7.2 Calculation of residual signals for a simplified downmix

[0101] The following sections discuss the calculation of residual signals and some considerations regarding the structure of channel signals in multi-channel audio signals.

[0102] In Unified Speech Acoustic Coding (USAC), there is no defined residual signal when so-called "simplified downmix" is used. Therefore, no partial waveform preservation coding is possible. However, the following describes a method for calculating the residual signal for so-called "simplified downmix".

[0103] Parametric upmix coefficient u d1 ,u d2 While the parameter bands are calculated separately, the "simplified downmix" weights d1 and d2 are calculated separately for each scale factor (coefficient) band. Therefore, the coefficient w used to calculate the residual signal is calculated separately. r1 ,w r2This cannot be calculated directly from spatial parameters (as is the case for classical MPEG surround), but it may require being determined for each scale factor (coefficient) band from the downmix coefficient and mixpmix coefficient.

[0104] Here, if L and R are input channels and D is the downmix channel, the residual signal res must satisfy the following characteristics. TIFF2026139655000002.tif23150

[0105] This is achieved by calculating the residuals as follows: TIFF2026139655000003.tif12150 Here, we use the following downmix weights. TIFF2026139655000004.tif28150

[0106] The residual upmix coefficient u used by the decoder r,1 ,u r,2 The method is preferably selected in a way that ensures robust decoding. Since the simplified downmix has asymmetric properties (in contrast to fixed-weighted MPEG surround), a spatial parameter-dependent upmix is ​​applied, for example, using the following upmix coefficients. TIFF2026139655000005.tif16150

[0107] Another option is to define residual upmix coefficients that are orthogonal to the upmix coefficients of the downmix signal, as follows: TIFF2026139655000006.tif17150

[0108] In other words, the audio decoder can obtain the downmix signal D using a linear combination of the left channel signal L (first channel signal) and the right channel signal R (second channel signal). Similarly, the residual signal res is obtained using a linear combination of the left channel signal L and the right channel signal R (or, generally, the first channel signal and the second channel signal of a multichannel audio signal).

[0109] For example, in equations (5) and (6), simplified downmix weights d1, d2 and parametric upmix coefficients u d,1 , u d,2 and residual upmix coefficients u r,1 , u r,2 are determined, it can be seen that downmix weights w for obtaining the residual signal res r,1 , w r,2 can be obtained. Furthermore, u r,1 , u r,2 can be derived from u d,1 , u d,2 using equations (7) and (8) or equation (9). The simplified downmix weights d1, d2, similarly to the parametric upmix coefficients u d,1 , u d,2 , can be obtained by a conventional method.

[0110] 7.3 Encoding Process

[0111] In the following, some details regarding the encoding process will be described. The encoding can be performed, for example, by the multichannel audio encoder 100, or by any other suitable means or computer program.

[0112] Preferably, the amount of transmitted residual is determined by a psychoacoustic model of an encoder (e.g., a multi-channel audio encoder), depending on the audio signal (e.g., the channel signals of a multi-channel audio signal 110) and the available bitrate. The transmitted residual signal can be used, for example, for partial waveform preservation or to avoid signal cancellation caused by the downmixing method used (e.g., the downmixing method described by equation (1) above).

[0113] 7.3.1 Partial waveform storage

[0114] The following describes how partial waveform preservation can be achieved. For example, the calculated residual (e.g., residual res by equation (4)) is transmitted in full band or band-limited to provide partial waveform preservation within the residual bandwidth. The residual portion, which is perceived as perceptually irrelevant by the psychoacoustic model, can be quantized to zero (e.g., based on the residual signal 126 when providing the encoded representation 112). This includes, but is not limited to, reducing the transmitted residual bandwidth at runtime (which can be thought of as changing the amount of residual signal included in the encoded representation). The system can also allow for bandpass-style erasure of the residual signal portion, since the lost signal energy is restored by a decoder (e.g., multi-channel audio decoder 200 or multi-channel audio decoder 300). Thus, while background noise can be parameterically encoded to reduce the residual bitrate, residual coding can, for example, be applied only to the tonal components of the signal while preserving their phase relationships. In other words, the residual signal 126 can be included in the encoded representation 112 only for frequency bands and / or time portions where the multichannel audio signal 110 (or at least one channel signal of the multichannel audio signal 110) is found to be tonal (e.g., by residual signal processing 130). In contrast, the residual signal 126 can be excluded from the encoded representation 112 for frequency bands and / or time portions where the multichannel audio signal 110 (or at least one or more channel signals of the multichannel audio signal 110) is identified as noise-like. Thus, the amount of residual signal included in the encoded representation varies according to the multichannel audio signal.

[0115] 7.3.2 Preventing signal cancellation in downmixing

[0116] The following describes how signal cancellation in downmixing can be prevented (or compensated for).

[0117] For low-bitrate applications, parametric coding (which relies primarily or exclusively on a parameter 124 describing the inter-channel dependencies of a multi-channel audio signal) is applied instead of waveform-preserving coding (which relies primarily on a residual signal 126 in addition to, for example, a downmix signal 122). Here, the residual signal 126 is used only to compensate for signal cancellation in the downmix 122 in order to minimize the bit usage of the residual. Unless signal cancellation is detected in the downmix 122, the system operates in parametric mode (on the audio decoder side) using a decorrelator. For example, when signal cancellation occurs for a fading tonal signal, the residual signal 126 is transmitted for the faulty signal portion (e.g., frequency band and / or time portion). Thus, the signal energy can be recovered by the decoder.

[0118] 7.4 Decryption Process

[0119] 7.4.1 Overview

[0120] In the decoder (e.g., multi-channel audio decoder 200 or multi-channel audio decoder 300), the transmitted downmix and residual signals (e.g., downmix signal 222 or residual signal 226) are decoded by the core decoder and fed to the MPEG surround decoder along with the decoded MPEG surround payload. The residual upmix coefficient for a classical MPS downmix is ​​invariant, and the residual upmix coefficient for a simplified downmix is ​​defined by equations (7), (8), and / or (9). In addition, the output of the uncorrelatedizer and its weighting coefficients are calculated with respect to parametric decoding. The residual signal and the output of the uncorrelatedizer are weighted, and both are mixed into the output signal. Thus, the weighting coefficients are determined by measuring the energy of the residual and uncorrelatedizer signals.

[0121] In other words, the residual upmix factor (or coefficient) can be determined by measuring the energy of the residual and uncorrelated signals.

[0122] For example, the downmix signal 222 is provided based on the encoded representation 210, and the uncorrelated signal 224 is generated based on parameters derived from the downmix signal 222 or contained in the encoded representation 210 (or otherwise). The residual upmix coefficient is a parametric upmix coefficient u, generated by the decoder according to, for example, equations (7) and (8). d,1 ,u d,2 It can be derived from the parametric upmix coefficient u d,1 ,u d,2 These can be obtained based on the encoded representation 210, for example, by directly deriving them from spatial data contained in the encoded representation 210 (e.g., from inter-channel correlation coefficients and inter-channel level difference coefficients, or from inter-object correlation coefficients and inter-object level differences).

[0123] Upmix coefficients for the uncorrelatedizer output (or output) can be obtained for conventional MPEG surround decoding. However, weighting coefficients for the weighting of the uncorrelatedizer output (or output(s)) can be determined based on the energy of the residual signal (and possibly also based on the energy of the uncorrelatedizer signal or signal(s)) so that the weights describing the contribution of the uncorrelated signal in the weighted combination are determined according to the residual signal.

[0124] 7.4.2 Exemplary Embodiments

[0125] Exemplary embodiments are described below with reference to Figure 7. However, it should be noted that the concepts described in this specification may also be applied to the multi-channel audio decoder 200 or 300 shown in Figures 2 and 3.

[0126] Figure 7 shows a schematic block diagram (or flow diagram) of a decoder (e.g., a multi-channel audio decoder). The decoder in Figure 7 is represented as 700. Decoder 700 is configured to receive a bitstream 710 and provide a first output channel signal 712 and a second output channel signal 714 based thereon. Decoder 700 includes a core decoder 720 configured to receive the bitstream 710 and provide a downmix signal 722, a residual signal 724, and spatial data 726 based thereon. For example, the core decoder 720 can provide a time-domain or transformation-domain representation (e.g., frequency-domain representation, MDCT-domain representation, QMF-domain representation) of the downmix signal represented by bitstream 710 as the downmix signal. Similarly, the core decoder 720 can provide a time-domain or transformation-domain representation of the residual signal 724 represented by bitstream 710. Furthermore, the core decoder 720 can provide one or more spatial parameters 726, such as one or more inter-channel correlation parameters, inter-channel level difference parameters, etc.

[0127] The decoder 700 also includes a decorrelator 730 configured to provide a decorrelated signal 732 based on the downmix signal 722. Any well-known decorrelation concept can be used by the decorrelator 730. Furthermore, the decoder 700 also receives spatial data 726 and upmix parameters (e.g., upmix parameter u dmx,1 ,u dmx,2 ,u dec,1 ,u dec,2 The decoder 700 further comprises an upmix coefficient calculator 740 configured to provide the upmix parameters 742 (also referred to as upmix coefficients) provided by the upmix coefficient calculator 740 based on spatial data 726. For example, the upmixer 750 obtains two upmixed versions 752, 754 of the downmix signal 722 by taking the upmix coefficients (e.g., u dmx,1 ,u dmx,2 The downmix signal 722 can be scaled using ). Furthermore, the upmixer 750 is also configured to apply one or more upmix parameters (e.g., two upmix parameters) to the uncorrelated signal 732 provided by the uncorrelatedizer 730 in order to obtain a first upmixed (scaled) version 756 and a second upmixed (scaled) version 758 of the uncorrelated signal 732. Furthermore, the upmixer 750 is also configured to apply one or more upmix coefficients (e.g., two upmix coefficients) to the residual signal 724 in order to obtain a first upmixed (scaled) version 760 and a second upmixed (scaled) version 762 of the residual signal 724.

[0128] The decoder 700 also includes a weight calculator 770 configured to measure the energies of upmixed (scaled) versions 756,758 of the uncorrelated signal 752 and the energies of upmixed (scaled) versions 760,762 of the residual signal 724. Furthermore, the weight calculator 770 is configured to provide one or more weight values ​​772 to a weighter 780. The weighter 780 is configured to use one or more weight values ​​772 provided by the weight calculator 770 to obtain a first upmixed (scaled), weighted version 782 of the uncorrelated signal 732, a second upmixed (scaled), weighted version 784 of the uncorrelated signal 732, a first upmixed (scaled), weighted version 786 of the residual signal 724, and a second upmixed (scaled), weighted version 788 of the residual signal 724. The decoder also includes a first adder 790 configured to sum a first upmixed (scaled) version 752 of the downmix signal 720, a first upmixed (scaled) and weighted version 782 of the uncorrelated signal 732, and a first upmixed (scaled) and weighted version 786 of the residual signal 724, in order to obtain a first output channel signal 712. Furthermore, the decoder includes a second adder 792 configured to sum a second upmixed version 754 of the downmix signal 720, a second upmixed (scaled) and weighted version 784 of the uncorrelated signal 732, and a second upmixed (scaled) and weighted version 788 of the residual signal 724, in order to obtain a second output channel signal 714.

[0129] However, it should be noted that the weighter 780 does not need to weight all signals 756, 758, 760, and 762. For example, in some embodiments, it may suffice to weight only signals 756 and 758, leaving signals 760 and 762 unaffected (practically, so that signals 760 and 762 are directly applied to adders 790 and 792). Alternatively, however, the weighting of residual signals 760 and 762 can be varied over time. For example, the residual signals can be faded in or faded out. For example, the weighting (or weighting coefficient) of the uncorrelated signals can be smoothed over time, and the residual signals can be faded in or faded out accordingly.

[0130] Furthermore, it should be noted that the weighting performed by the weighter 780 and the upmixing applied by the upmixer 750 can be performed as a combined operation, and the weight calculation can be performed directly using the uncorrelated signal 732 and the residual signal 724.

[0131] The following provides some further details regarding the functionality of the decoder 700.

[0132] The combined residual and parametric coding modes can be signaled, for example, in a quasi-backward compatible manner, by signaling the residual bandwidth of one parameter band in the bitstream. Thus, the legacy decoder still passes through and decodes the bitstream by switching to parametric decoding on the first parameter band. The legacy bitstream with residual bandwidth does not contain residual energy on the first parameter band and becomes parametrically decoded in the proposed novel decoder.

[0133] However, within a 3D audio codec system, combined residual and parametric coding can be used in combination with other core decoder tools such as quad-channel elements, allowing the decoder to explicitly detect legacy bitstreams and enable their decoding in the normal band-limited residual coding mode. The actual residual bandwidth is determined by the decoder at runtime and is therefore preferably not explicitly signaled. The calculation of the upmix coefficients is set to parametric mode instead of residual coding mode. The energy Edec of the weighted uncorrelatedizer output and the energy of the weighted residual signal Eres are calculated for each upmix channel ch for each frame and the hybrid band hb across all time slots ts, as follows: TIFF2026139655000007.tif26150

[0134] TIFF2026139655000008.tif34129

[0135] The residual signals (e.g., upmixed residual signal 760 or upmixed residual signal 762) are added to the output channels (e.g., output channels 712, 714) with a weight of 1. The uncorrelated signals (e.g., upmixed uncorrelated signal 756 or upmixed uncorrelated signal 758) can be weighted by a coefficient r calculated as follows (e.g., by the weighter 780). TIFF2026139655000009.tif24154 Here, E dec (hb) is the uncorrelated signal x for the frequency band hb. dec This represents the weighted energy value of E res (hb) is the residual signal x for the frequency band hb. res This represents the weighted energy value.

[0136] If the residual (for example, residual signal 724) is not transmitted, for example, E resIf = 0, then r (a coefficient that can be applied by the weighter 780 and can be considered as the weighted value 772) becomes 1, which is purely parametric decoding. If the residual energy (e.g., the energy of the upmixed residual signal 760 and / or upmixed residual signal 762) exceeds the uncorrelatedizer energy (e.g., the energy of the upmixed uncorrelated signal 756 or upmixed uncorrelated signal 758), then, for example, E res > E dec In this case, the coefficient r can be set to zero, thus disabling the decorrelator and enabling partial waveform preservation decoding (which can be considered residual coding). In the upmixing process, both the weighted decorrelator output (e.g., signals 782,784) and the residual signal (e.g., signals 786,788 or signals 760,762) are added to the output channel (e.g., signals 712,714).

[0137] In conclusion, this becomes a matrix-style upmix rule. TIFF2026139655000010.tif16160 Here, ch1 represents one or more time-domain samples or transformed-domain samples of the first output audio signal, ch2 represents one or more time-domain samples or transformed-domain samples of the second output audio signal, x dmx x represents one or more time-domain samples or transformed-domain samples of the downmix signal, dec x represents one or more time-domain samples or transformed-domain samples of the uncorrelated signal, res represents one or more time-domain samples or transform-domain samples of the residual signal, u dmx,1 represents the downmix signal upmix parameter for the first output audio signal, and u dmx,2 represents the downmix signal upmix parameter for the second output audio signal, and u dec,1 represents the uncorrelated signal upmix parameter for the first output audio signal, and u dec,2represents the uncorrelated signal upmix parameter for the second output audio signal, max represents the maximum operator, and r represents a coefficient that describes the weighting of the uncorrelated signal according to the residual signal.

[0138] Upmix coefficient U dmx,1 ,U dmx,2 ,U dec,1 ,U dec,2 This is calculated with respect to the MPS2-1-2 parametric mode. For details, refer to the MPEG surround concept standard mentioned above.

[0139] In summary, embodiments of the present invention construct a concept that provides an output channel signal based on a downmix signal, a residual signal, and spatial data, with flexible adjustment of the weighting of the uncorrelated signals without any significant signaling overhead.

[0140] 7.5 Modified Embodiments

[0141] While several embodiments have been described in the context of apparatus, it is clear that these embodiments also represent descriptions of corresponding methods, where blocks or devices correspond to method steps or features of method steps. Similarly, embodiments described in the context of method steps also represent descriptions of corresponding blocks, items, or features of the corresponding apparatus. Some or all method steps can be performed by (or using) hardware devices, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such devices.

[0142] The encoded audio signal of the present invention can be stored in a digital storage medium or transmitted over a transmission medium (e.g., a wireless transmission medium or a wired transmission medium (e.g., the Internet)).

[0143] Depending on the specific implementation requirements, embodiments of the present invention may be implemented in hardware or in software. Implementation may be carried out using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, which has electronically readable control signals stored thereon and cooperates (or can cooperate) with a computer system programmable to perform each method. Therefore, the digital storage medium may be computer-readable.

[0144] Some embodiments of the present invention include a data carrier having an electronically readable control signal and capable of cooperating with a programmable computer system on which one of the methods described in the present specification is performed.

[0145] Generally, embodiments of the present invention can be implemented as a computer program product having program code that is operable to perform one of the methods when the computer program product runs on a computer. The program code can be stored, for example, on a machine-readable carrier.

[0146] Other embodiments include a computer program stored in a machine-readable carrier for performing one of the methods described in this specification.

[0147] In other words, an embodiment of the method of the present invention is a computer program having program code for executing one of the methods described in the present specification when the computer program is running on a computer.

[0148] A further embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) having a computer program recorded thereon that performs one of the methods described in the present specification. The data carrier, digital storage medium or recording medium is generally tangible and / or non-transient.

[0149] A further embodiment of the method of the invention is, therefore, a data stream or sequence of signals representing a computer program that performs one of the methods described in this specification. The data stream or sequence of signals may be configured to be transmitted, for example, over a data communication connection, such as the Internet.

[0150] Further embodiments include processing means, such as a computer or a programmable logic device, configured or adapted to perform one of the methods described in this specification.

[0151] A further embodiment comprises a computer on which a computer program for performing one of the methods described in this specification is installed.

[0152] Further embodiments of the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program that performs one of the methods described in the present specification to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server that transfers the computer program to the receiver.

[0153] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field-programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0154] The embodiments described above are merely illustrative of the principles of the present invention. Modifications and changes to the configurations and details described in this specification will be understood to be obvious to those skilled in the art. The present invention is therefore intended to be limited only by the immediate claims and not by the specific details provided by the descriptions and explanations of the embodiments in this specification.

[0155] 7.6 Further Embodiments

[0156] In the following, other embodiments of the present invention will be described with reference to Figure 8, which shows a schematic block diagram of a so-called hybrid residual decoder.

[0157] The hybrid residual decoder 800 shown in Figure 8 is very similar to the decoder 700 shown in Figure 7, and the above explanation is referred to. However, in the hybrid residual decoder 800, additional weighting (in addition to the application of upmix parameters) is applied only to the upmixed uncorrelated signal (which corresponds to signals 756 and 758 in decoder 700), and not to the upmixed residual signal (which corresponds to signals 760 and 762 in decoder 700). Thus, the weighting of the hybrid residual decoder 800 is somewhat simpler than the weighting in decoder 700, but it agrees well with the weighting by equation (14), for example.

[0158] The combined parametric and residual decoding (hybrid residual coding) shown in Figure 8 will be explained in some more detail below.

[0159] However, an overview is provided first.

[0160] In addition to using either a decorrelator-based mono-to-stereo upmix or residual coding as described in ISO / IEC 23003-3 (Section 7.11.1), hybrid residual coding allows for the dependent coupling of signals in both modes. As illustrated in Figure 8, the residual signal and the decorrelator output are mixed using time- and frequency-dependent weighting coefficients depending on the signal energy and spatial parameters.

[0161] The decryption process is described below.

[0162] TIFF2026139655000011.tif62170

[0163] The upmixing process is divided into the downmix, the uncorrelated output, and the residual. dmx It is calculated using the following formula. TIFF2026139655000012.tif19170

[0164] Upmixed uncorrelated output u dec It is calculated using the following formula. TIFF2026139655000013.tif16158

[0165] Upmixed residual signal u res It is calculated using the following formula. TIFF2026139655000014.tif17158

[0166] Energy E of the upmixed residual signal resThe energy Edec of the upmixed uncorrelated output is calculated as a sum over both the output channel ch and all time slots ts of one frame for each hybrid band, as follows: TIFF2026139655000015.tif24157

[0167] The upmixed uncorrelated output is obtained by applying weighting coefficients r calculated per frame for each hybrid band, as follows: dec It is weighted using [this method]. TIFF2026139655000016.tif44158 Here, ε is a small number to prevent division by zero (e.g., ε = 1e-9 or 0 < ε <= 1e-5). However, in some embodiments, ε is set to zero ("E res <ε> to "E res It can be replaced with "=0".

[0168] All three upmix signals are added together to form the decoded output signal.

[0169] 8. Conclusion

[0170] In conclusion, embodiments of the present invention construct a combined residual and a parametric encoding.

[0171] The present invention constructs a method for signal-dependent coupling of parametric and residual coding for joint stereo coding based on the USAC Integrated Stereo Tool. Instead of using a fixed residual bandwidth, the amount of transmitted residual is signal-dependently determined by encoder, time, and frequency variability. On the decoder side, the required amount of decorrelation between output channels is generated by mixing the residual signal with the decorrelationizer output. Thus, the corresponding audio coding / decoding system can mix between full parametric coding and waveform-preserving residual coding at runtime, depending on the coded signal.

[0172] Embodiments relating to the present invention are superior to conventional solutions. For example, in USAC, the MPEG Surround 2-1-2 system is used for parametric stereo coding or integrated stereo, transmitting a band-limited or full-bandwidth residual signal for partial waveform preservation. When a band-limited residual is transmitted, a parametric upmix is ​​applied to the residual bandwidth using a decorrelator. The drawback of this method is that the residual bandwidth is set to a fixed value during encoder initialization.

[0173] In contrast, embodiments of the present invention enable signal dependency adaptation of residual bandwidth or switching to parametric coding. Furthermore, when the downmixing process in parametric coding mode results in signal cancellation for poor phase relationships, embodiments of the present invention enable the restoration of lost signal portions (e.g., by providing appropriate residual signals). It should be noted that the simplified downmixing method results in less signal cancellation for parametric coding than the classical MPS downmixing method. However, conventional simplified downmixing cannot be used for partial waveform preservation because residual signals are not defined in USAC, whereas embodiments of the present invention enable waveform restoration (e.g., selective partial waveform restoration for signal portions where partial waveform restoration appears important).

[0174] In conclusion, embodiments of the present invention construct an apparatus, method, or computer program for audio encoding or decoding as described in the present specification.

Claims

1. A multichannel audio decoder (200;300;700;800) that provides at least two output audio signals (212,214;312,314;712,714) based on an encoded representation (210;310;710), The multi-channel audio decoder is configured to perform a weighted combination (220; 780; 790; 792) of a downmix signal (222; 752; 754), an uncorrelated signal (224; 756; 758), and a residual signal (226; 760; 762; res) in order to obtain one of the output audio signals (212; 214; 712; 714), The multichannel audio decoder uses weights (232;r;r) according to the residual signal to describe the contribution of the uncorrelated signal in the weighted coupling. dec ) configured to determine Multi-channel audio decoder.

2. The multichannel audio decoder according to claim 1, wherein the multichannel audio decoder is configured to determine weights that describe the contribution of the uncorrelated signal in the weighted coupling, in accordance with the uncorrelated signal.

3. The multi-channel audio decoder, based on the encoded representation, upmixes the parameters (u dmx,1 , u dmx,2 , u dec,1 , u dec,2 , u r,1 , u r,2 It is configured to obtain weights (232;r;r) that describe the contribution of the uncorrelated signal in the weighted combination according to the upmix parameter. dec A multichannel audio decoder according to claim 1 or 2, configured to determine ).

4. The multi-channel audio decoder is configured to determine a weight (232; r; r dec ) describing the contribution of the decorrelated signal in said weighted combination such that the weight of the decorrelated signal decreases as the energy of the residual signal increases, the multi-channel audio decoder according to any one of claims 1 to 3.

5. The multi-channel audio decoder, when the energy of the residual signal is zero, sets the uncorrelated signal upmix parameter (u dec,1 , u dec,2 ;u dec (hb,ts,ch);u dec The maximum weight determined by (ch, ts) is related to the uncorrelated signal, and the residual signal weight coefficient (u r,1 , u r,2 ;u res (hb,ts,ch);u res If the energy of the residual signal weighted by (ch, ts) is greater than or equal to the energy of the uncorrelated signal weighted by the uncorrelated signal upmix parameter, then the zero weight is associated with the uncorrelated signal, and the weights describing the contribution of the uncorrelated signal in the weighted combination are (232; r; r dec A multichannel audio decoder according to any one of claims 1 to 4, configured to determine ).

6. The multi-channel audio decoder calculates a factor (r, r) according to the weighted energy value of the uncorrelated signal and the weighted energy value of the residual signal. dec To determine the factor and obtain a weight describing the contribution of the uncorrelated signal to one of the output audio signals based on the factor, or to use the factor as a weight describing the contribution of the uncorrelated signal to one of the output audio signals, the weighted energy value (E) of the uncorrelated signal is weighted according to one or more uncorrelated signal upmix parameters. dec (hb); E dec ) is calculated, and the weighted energy value (E) of the residual signal is weighted using one or more residual signal upmix parameters. res (hb); E res A multichannel audio decoder according to any one of claims 1 to 5, configured to calculate ()

7. The multichannel audio decoder obtains the coefficient (r) to the uncorrelated signal upmix parameter (u) in order to obtain the weight that describes the contribution of the uncorrelated signal to one of the output audio signals. dec,1 , u dec,2 ;u dec (hb,ts,ch);u dec A multichannel audio decoder according to claim 6, configured to multiply by (ch, ts).

8. The multi-channel audio decoder uses the weighted energy value (E) of the uncorrelated signal. dec (hb); E dec A multichannel audio decoder according to claim 6 or 7, configured to calculate the energy of the uncorrelated signal weighted using uncorrelated signal upmix parameters across a plurality of upmix channels (ch) and time slots (ts) in order to obtain the following:

9. The multi-channel audio decoder uses the weighted energy value (E) of the residual signal. res (hb); E res A multichannel audio decoder according to any one of claims 6 to 8, configured to calculate the energy of the residual signal weighted with residual signal upmix parameters over a plurality of upmix channels (ch) and time slots (ts) in order to obtain the following:

10. The multi-channel audio decoder uses the weighted energy value (E) of the uncorrelated signal. dec (hb); E dec ) and the weighted energy value (E) of the residual signal res (hb); E res The coefficient (r; r) is determined according to the difference between the two. dec A multichannel audio decoder according to any one of claims 6 to 9, configured to calculate ).

11. The aforementioned multi-channel audio decoder is - The difference between the weighted energy value of the uncorrelated signal and the weighted energy value of the residual signal, - The weighted energy value of the uncorrelated signal and The coefficient (r; r) is determined according to the ratio of the coefficients (r; r) dec A multichannel audio decoder according to claim 10, configured to calculate ).

12. The multi-channel audio decoder is configured to determine weights that describe the contribution of the uncorrelated signal to two or more output audio signals. The multi-channel audio decoder uses the weighted energy value (E) of the uncorrelated signal. dec (hb); E dec ) and the uncorrelated signal upmix parameter of the first channel (u dec,1 Based on this, the contribution of the uncorrelated signal to the first output audio signal is determined, The multi-channel audio decoder uses the weighted energy value (E) of the uncorrelated signal. dec (hb); E dec ) and the uncorrelated signal upmix parameter of the second channel (u dec,2 Based on this, it is configured to determine the contribution of the uncorrelated signal to the second output audio channel, A multichannel audio decoder according to any one of claims 6 to 11.

13. The multi-channel audio decoder has residual energy (E res (hb); E res ) is the energy of the uncorrelatedizer (E de c(hb); E dec A multichannel audio decoder according to any one of claims 1 to 12, configured to disable the contribution of the uncorrelated signal to the weighted coupling when it exceeds ).

14. The aforementioned multi-channel audio decoder is It is configured to calculate two output audio signals, ch1 and ch2, Here, ch1 represents one or more time-domain samples or transformed-domain samples of the first output audio signal, ch2 represents one or more time-domain samples or transformed-domain samples of the second output audio signal, x dmx x represents one or more time-domain samples or transformation-domain samples of the downmix signal, dec x represents one or more time-domain samples or transformed-domain samples of the uncorrelated signal, res represents one or more time-domain samples or transformed-domain samples of the residual signal, u dmx,1 represents the downmix signal upmix parameter for the first output audio signal, u dmx,2 represents the downmix signal upmix parameter for the second output audio signal, u dec,1 represents the uncorrelated signal upmix parameter for the first output audio signal, u dec,2 represents the uncorrelated signal upmix parameter for the second output audio signal, max represents the maximum operator, and r represents a coefficient that describes the weighting of the uncorrelated signal according to the residual signal. A multichannel audio decoder according to any one of claims 1 to 13.

15. The aforementioned multi-channel audio decoder is The coefficient r is configured to be calculated accordingly. Here, E dec (hb) or E dec The uncorrelated signal x for the frequency band hb is dec This represents the weighted energy value of E res (hb) or E res The residual signal x with respect to the frequency band hb is res Represents the weighted energy values, The multi-channel audio decoder according to claim 14.

16. The aforementioned multi-channel audio decoder is The weighted energy values ​​of the residual signal are configured to be calculated accordingly. Here, res x represents the residual signal upmix parameter for the frequency band hb, time slot ts, and upmix channel ch, res This represents a time-domain sample or transformed-domain sample of the uncorrelated signal for the frequency band hb, time slot ts, and upmix channel ch. The multi-channel audio decoder according to claim 15.

17. The audio decoder determines the weights (232;r;r) that describe the contribution of the uncorrelated signal in the weighted coupling, according to the band-by-band determination of the weighted energy values ​​of the residual signal. dec A multichannel audio decoder according to any one of claims 1 to 16, configured to determine the frequency for each band.

18. The audio decoder according to any one of claims 1 to 17, wherein the audio decoder is configured to determine weights for each frame of the output audio signal that describe the contribution of the uncorrelated signal in the weighted coupling.

19. The audio decoder according to any one of claims 1 to 18, wherein the multichannel audio decoder is configured to variably adjust the weights that describe the contribution of the residual signal in the weighted coupling.

20. A multichannel audio decoder (200;300;700;800) that provides at least two output audio signals (212,214;312,314;712,714) based on an encoded representation (210;310;710), The multi-channel audio decoder is configured to acquire one of the output audio signals based on an encoded representation of the downmix signal (222; 722), a plurality of encoded spatial parameters (726), and an encoded representation of the residual signal (226; 724). The multichannel audio decoder is configured to mix between parametric coding and residual coding according to the residual signal.

21. A multichannel audio encoder (100) that provides an encoded representation (112) of a multichannel audio signal (110), The multi-channel audio encoder is configured to acquire a downmix signal (122) based on the multi-channel audio signal, provide parameters (124) describing the inter-channel dependencies of the multi-channel audio signal, and provide a residual signal (126). The multi-channel audio encoder is configured to change the amount of residual signal included in the encoded representation according to the multi-channel audio signal. Multi-channel audio encoder.

22. The multichannel audio encoder according to claim 21, wherein the multichannel audio encoder is configured to change the bandwidth of the residual signal according to the multichannel audio signal.

23. The multichannel audio encoder according to claim 21 or 22, wherein the multichannel audio encoder is configured to select frequency bands in which the residual signal is included in the encoded representation according to the multichannel audio signal.

24. The multichannel audio encoder according to claim 23, wherein the multichannel audio encoder is configured to selectively include the residual signal in the encoded representation for frequency bands in which the multichannel audio signal is tonal.

25. The multichannel audio encoder according to any one of claims 21 to 24, wherein the multichannel audio encoder is configured to selectively include the residual signal in the encoded representation for time portions and / or frequency bands in which the formation of the downmix signal results in the cancellation of signal components of the multichannel audio signal.

26. The multichannel audio encoder according to claim 25, wherein the multichannel audio encoder is configured to detect the cancellation of signal components of the multichannel audio signal in the downmix signal, and the multichannel audio encoder is configured to activate the provision of the residual signal in response to the result of the detection.

27. The multichannel audio encoder according to any one of claims 21 to 26, wherein the multichannel audio encoder is configured to calculate the residual signal using a linear combination of at least two channel signals of the multichannel audio signal, according to an upmix coefficient used on the multichannel decoder side.

28. The multi-channel audio encoder determines and encodes the upmix coefficient. Alternatively, the upmix coefficient is derived from parameters describing the inter-channel dependency of the multi-channel audio signal. A multichannel audio encoder according to claim 27, configured as described above.

29. The multichannel audio encoder according to any one of claims 21 to 28, wherein the multichannel audio encoder is configured to determine the amount of the residual signal included in the encoded representation as a time variable using a psychoacoustic model.

30. The multichannel audio encoder according to any one of claims 21 to 29, wherein the multichannel audio encoder is configured to determine the amount of the residual signal included in the encoded representation as a time variable according to the currently available bitrate.

31. A method (500) for providing at least two output audio signals based on an encoded representation, The step (520) of performing a weighted combination of the downmix signal, the uncorrelated signal and the residual signal in order to obtain one of the output audio signals, The weights describing the contribution of the uncorrelated signal in the weighted combination are determined according to the residual signal (510). method.

32. A method (600) for providing at least two output audio signals based on an encoded representation, The process includes the step (610) of obtaining one of the output audio signals based on the encoded representation of the downmix signal, a plurality of encoding space parameters, and the encoded representation of the residual signal, Mixing is performed between parametric coding and residual coding according to the residual signal (620). method.

33. A method (400) for providing an encoded representation of a multichannel audio signal, The steps include: (410) acquiring a downmix signal based on the multi-channel audio signal; The steps include (420) providing parameters that describe the inter-channel dependencies of the multi-channel audio signal, The process includes the step (430) of providing a residual signal, The amount of residual signal included in the encoded representation is changed according to the multi-channel audio signal (440). method.

34. A computer program that, when running on a computer, performs the method according to any one of claims 31 to 33.

Citation Information

Patent Citations

  • WOIEC13818-7

  • WOIEC23003-2

  • WOIEC23003-3

  • WOIEC23003-1