Multichannel audio decoder, multichannel audio encoder, method, and computer program using residual signal-based adjustment of the contribution of uncorrelated signal
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
- Filing Date
- 2023-04-21
- Publication Date
- 2026-05-19
AI Technical Summary
Existing multichannel audio encoding and decoding technologies lack efficient methods for reconstructing high-quality audio signals, particularly in scenarios where residual signals are insufficient or not adequately utilized.
A multichannel audio decoder that performs a weighted combination of a downmix signal, an uncorrelated signal, and a residual signal, adjusting weights based on residual signal energy to enhance reconstruction quality without additional signaling, and a multichannel audio encoder that varies residual signal inclusion based on audio characteristics.
Enables high-quality reconstruction of multichannel audio signals with flexible adaptation to signal characteristics, achieving efficient bitrate usage and improved audio quality by selectively using residual signals.
Smart Images

Figure 0007862341000016 
Figure 0007862341000017 
Figure 0007862341000018
Abstract
Description
[Technical Field]
[0001] Embodiments of the present invention relate to a multi-channel audio decoder that provides at least two output audio signals based on an encoded representation.
[0002] Another embodiment of the present invention relates to a multichannel audio encoder that provides an encoded representation of a multichannel audio signal.
[0003] Another embodiment of the present invention relates to a method for providing at least two output audio signals based on an encoded representation.
[0004] Another embodiment of the present invention relates to a method for providing an encoded representation of a multichannel audio signal.
[0005] Another embodiment of the present invention relates to a computer program for performing one of the above methods.
[0006] generally Several embodiments of the present invention are , remainder difference coding and parametric Combined coding Regarding. [Background technology]
[0007] In recent years, the demand for storing and transmitting audio content has steadily increased. Furthermore, the demand for quality in storing and transmitting audio content has also steadily increased. Therefore, concepts for encoding and decoding audio content have been strengthened. For example, the so-called "Advanced Audio Coding" (AAC), described in Non-Patent Document 1, has been developed.
[0008] Furthermore, for example, several spatial extensions such as the so-called "MPEG Surround" concept described in Non-Patent Document 2 have been constructed. Additionally, Non-Patent Document 3 describes additional improvements to the encoding and decoding of spatial information of audio signals regarding so-called spatial audio object coding. Furthermore, a flexible (switchable) audio encoding / decoding concept that describes the so-called "integrated speech and audio coding" concept, encodes both general audio signals and speech signals with good coding efficiency, and provides the possibility of handling multi-channel audio signals is defined in Non-Patent Document 4.
Prior Art Documents
Non-Patent Documents
[0009]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Non-Patent Document 4
Summary of the Invention
Problems to be Solved by the Invention
[0010] However, it is desirable to provide even more advanced concepts for the efficient encoding and decoding of multi-channel audio signals.
Means for Solving the Problems
[0011] Embodiments of the present invention construct a multichannel audio decoder that provides at least two output audio signals based on an encoded representation. The multichannel audio decoder is configured to obtain one of the output audio signals by performing a weighted combination of a downmix signal, an uncorrelated signal, and a residual signal. The multichannel audio decoder is configured to determine weights that describe the contribution of the uncorrelated signal in the weighted combination according to the residual signal.
[0012] In this embodiment of the present invention, the weights describing the contribution of the uncorrelated signal to the weighted combination of the downmix signal, the uncorrelated signal, and the residual signal are adjusted according to the residual signal. If so , output audio signal teeth Based on the encoded representation extremely efficient to It can be obtained Knowledge This is based on the following: Therefore, the weights describing the contribution of the uncorrelated signal in the weighted combination are applied to the residual signal. Dependent By adjusting, additional Parametric without transmitting control information coding (or mainly parametric) coding ) and residuals coding (or mostly residuals) coding ) between blend It is possible to (blend) (or fade). Furthermore, The residual signals included in the encoded representation proved to be excellent indicators of the weights describing the contribution of the uncorrelated signals in the weighted combination. Typically, The residual signal is (comparison target) Weak (or desired energy) of Restoration To (Insufficient) teeth Uncorrelated signal to ( comparison (Likely) Large Weighted, the residual signal (compared target) strong (or sufficient Desired energy of Restoration It is possible ) in the case teeth Uncorrelated signal to ( comparison target) Small Ku Adding weights is good Mashii Therefore Therefore, as mentioned above By concept , ( For example, desired energy characteristics and / or correlation characteristics are signaled by parameters and then restored by adding an uncorrelated signal. Parametric coding and , (Da Based on the UNMIX signal and come out Power audio signal ~ In some cases, the waveform of the output audio signal Even too ~ To restore The residual signal (used) Residual coding and of Gradual transition between This becomes possible Therefore, additional Signalin Guo Without having a bar head 、 Restoration Technology and further The quality of the restoration Decoded signal It is possible to adapt it.
[0013] In a preferred embodiment, the multichannel audio decoder is configured to determine weights describing the contribution of the uncorrelated signal in the weighted combination, according to the uncorrelated signal. By determining weights describing the contribution of the uncorrelated signal in the weighted combination, according to both the residual signal and the uncorrelated signal, the weights can be appropriately adjusted to the signal characteristics so as to achieve good quality restoration of at least two output audio signals based on the encoded representation (in particular, based on the downmix signal, the uncorrelated signal, and the residual signal).
[0014] In a preferred embodiment, the multichannel audio decoder is configured to obtain upmix parameters based on the encoded representation and to determine weights that describe the contribution of the uncorrelated signals in the weighted combination according to the upmix parameters. By considering the upmix parameters, it is possible to reconstruct desired characteristics of the output audio signals (such as desired correlation between output audio signals and / or desired energy characteristics of the output audio signals) and take desired values.
[0015] In a preferred embodiment, the multichannel audio decoder is configured to determine weights describing the contribution of the uncorrelated signal in the weighted coupling such that the weight of the uncorrelated signal decreases with increasing energy of one or more residual signals. This mechanism allows for adjusting the accuracy of the reconstruction of at least two output audio signals according to the energy of the residual signals. When the energy of the residual signals is relatively high, the weight of the contribution of the uncorrelated signal is relatively small so that the uncorrelated signal does not have a detrimental effect on the high quality of reproduction that would result from using the residual signals. In contrast, when the energy of the residual signals is relatively low or zero, a high weight is given to the uncorrelated signal so that the uncorrelated signal can efficiently bring the characteristics of the output audio signals to the desired value.
[0016] In a preferred embodiment, the multichannel audio decoder is configured to determine weights describing the contribution of the uncorrelated signal in the weighted combination such that when the energy of the residual signal is zero, the maximum weight determined by the uncorrelated signal upmix parameter is associated with the uncorrelated signal, and when the energy of the residual signal weighted using the residual signal weight coefficient is greater than or equal to the energy of the uncorrelated signal weighted by the uncorrelated signal upmix parameter, the zero weight is associated with the uncorrelated signal. This embodiment is based on the finding that the desired energy to be added to the downmix signal is determined by the energy of the uncorrelated signal weighted by the uncorrelated signal upmix parameter. Thus, it can be concluded that the uncorrelated signal no longer needs to be added when the energy of the residual signal weighted by the residual signal weight coefficient is greater than or equal to the energy of the uncorrelated signal weighted by the uncorrelated signal upmix parameter. In other words, when it is determined that the residual signal has sufficient energy (e.g., sufficient to reach a sufficient total energy), the uncorrelated signal is no longer used to provide at least two output audio signals.
[0017] In a preferred embodiment, the multichannel audio decoder is configured to calculate a weighted energy value of the uncorrelated signal weighted according to one or more uncorrelated signal upmix parameters, and to calculate a weighted energy value of the residual signal weighted using one or more residual signal upmix parameters (which may be equal to the residual signal weight coefficients described above), to determine a factor according to the weighted energy values of the uncorrelated signal and the weighted energy values of the residual signal, and to obtain a weight that describes the contribution of the uncorrelated signal to (at least) one audio output signal based on that factor. This procedure has been found to be well suited to the efficient calculation of weights that describe the contribution of the uncorrelated signal to one or more output audio signals.
[0018] In a preferred embodiment, the multichannel audio decoder multiplies the elements by an uncorrelated signal upmix parameter. and come out Power audio signal (At least) one of It is configured to obtain weights that describe the contribution of the uncorrelated signal to the given value. Using such a procedure By doing To determine the weights that describe the contribution of the uncorrelated signal in the weighted combination, (Described by the uncorrelated signal upmix parameter) One or more parameters that describe the desired signal characteristics of at least two output audio signals Ta and The relationship between the energy of the uncorrelated signal and the energy of the residual signal. and It is possible to consider both. Nina ru. In this way , (Reflected by the uncorrelated signal upmix parameter) Desired characteristics of the output audio signal Sex Still considering, parametric coding (or mainly parametric) coding ) and residuals coding (or mainly residuals) coding ) between blend It could either do (or fade).
[0019] In a preferred embodiment, the multichannel audio decoder is configured to obtain a weighted energy value for the uncorrelated signal by calculating the energy of the uncorrelated signal weighted using the uncorrelated signal upmix parameter across multiple upmix channels and time slots. Thus, it is possible to avoid large changes in the weighted energy value of the uncorrelated signal. Consequently, stable tuning of the multichannel audio decoder is achieved.
[0020] Similarly, a multi-channel audio decoder is configured to obtain a weighted energy value for the residual signal by calculating the energy of the residual signal weighted using the residual signal upmix parameters across multiple upmix channels and time slots. Thus, stable tuning of the multi-channel audio decoder is achieved, as large changes in the weighted energy value of the residual signal are avoided. However, the averaging period can be selected to be sufficiently short to allow for dynamic adjustment of the weighting.
[0021] In a preferred embodiment, the multichannel audio decoder is configured to compute elements according to the difference between the weighted energy values of the uncorrelated signal and the weighted energy values of the residual signal. The calculation comparing the weighted energy values of the uncorrelated signal and the weighted energy values of the residual signal allows for the supplementation of the residual signal (or a weighted version of the residual signal) with the uncorrelated signal (or a weighted version of the residual signal), and the weights are adjusted to describe the contribution of the uncorrelated signal to the need to provide at least two audio channel signals.
[0022] In a preferred embodiment, the multichannel audio decoder is configured to calculate elements according to the ratio between the difference between the weighted energy value of the uncorrelated signal and the weighted energy value of the residual signal, and the weighted energy value of the uncorrelated signal. Calculation of elements according to this ratio has been shown to yield long, characteristically good results. Furthermore, it should be noted that this ratio describes which portion of the total energy of the uncorrelated signal (weighted using the uncorrelated signal upmix parameter) is needed in the presence of the residual signal to achieve a good auditory impression (or, equivalently, to have substantially the same signal energy in the output audio signal compared to the case without the residual signal).
[0023] In a preferred embodiment, the multichannel audio decoder is configured to determine weights describing the contribution of the uncorrelated signal to two or more output audio signals. In this case, the multichannel audio decoder is configured to determine the contribution of the uncorrelated signal to a first output audio signal based on the weighting energy value of the uncorrelated signal and the uncorrelated signal upmix parameter of the first channel. Furthermore, the multichannel audio decoder is configured to determine the contribution of the uncorrelated signal to a second output audio channel based on the weighting energy value of the uncorrelated signal and the uncorrelated signal upmix parameter of the second channel. Thus, the two output audio signals can be delivered with reasonable effort and good audio quality, and the difference between the two output audio signals is taken into account by using the uncorrelated signal upmix parameter of the first channel and the uncorrelated signal upmix parameter of the second channel.
[0024] In a preferred embodiment, the multichannel audio decoder is configured to disable the contribution of the uncorrelated signal to weighted coupling when the residual energy exceeds the energy of the uncorrelatedizer (i.e., the energy of the uncorrelated signal or its weighted version). Thus, if the residual signal has sufficient energy and the residual energy exceeds the energy of the uncorrelatedizer, it is possible to switch to pure residual coding without using the uncorrelated signal.
[0025] In a preferred embodiment, the audio decoder determines the band-by-band weighted energy values of the residual signal. Dependent The system is configured to determine, for each band, the weights that describe the contribution of the uncorrelated signal in the weighted coupling. Therefore, additional Signalin Guo Without a bar head , small Even without that, the improvement of the two output audio signals In which frequency band parametric coding Should it be based on (or should it be primarily based on) ) , and Improvements to at least two output audio signals In which frequency band residual coding Should it be based on (or should it be primarily based on) )of It is possible to make decisions flexibly. In this way While keeping the weights of the uncorrelated signals relatively small, the residuals in all frequency bands coding It allows for a flexible decision on whether to perform waveform reconstruction (or at least partial waveform reconstruction) using (at least primarily) this method. Therefore , (Based primarily on the supply of uncorrelated signals) parametric coding and, (Based primarily on the supply of residual signals) residual coding and Select By selectively applying it, it is possible to obtain good audio quality.
[0026] In a preferred embodiment, the audio decoder is configured to determine weights for each frame of the output audio signal that describe the contribution of the uncorrelated signal in the weighted coupling. Thus, fine timing resolution can be obtained. Subsequent Parametric between frames coding (or mainly parametric) coding ) and residuals coding (or mainly residuals) coding ) flexible between Switching Possible Nina Therefore, With good temporal resolution, Regarding the characteristics of the audio signal te O Audio decoding can be adjusted.
[0027] Another embodiment of the present invention is based on an encoded representation. little To build a multi-channel audio decoder that provides at least two output audio signals, One of the output audio signals (at least one of them) Based on the encoded representation of the downmix signal, multiple encoded spatial parameters, and the encoded representation of the residual signal Take It is configured to obtain a multi-channel audio decoder. is remaining Difference signal Dependent ,parametric coding and residuals coding Between blend It is configured to do so. Therefore, additional Signalin Guo Without overhead, the best decoding mode (parametric) coding • Decoding vs. Residual coding A highly flexible audio decoding concept is achieved, allowing for the selection of decoding methods. Furthermore, the considerations mentioned above also apply.
[0028] Embodiments of the present invention construct a multichannel audio encoder that provides an encoded representation of a multichannel audio signal. The multichannel audio encoder is configured to acquire a downmix signal based on the multichannel audio signal. Furthermore, the multichannel audio encoder is configured to provide parameters that describe the inter-channel dependencies of the multichannel audio signal and to provide residual signals. Furthermore, the multichannel audio encoder is configured to vary the amount of residual signal included in the encoded representation according to the multichannel audio signal. By varying the amount of residual signal included in the encoded representation, the encoding process can be flexibly adjusted to the characteristics of the signal. For example, it is possible to include a relatively large amount of residual signal in the encoded representation for parts of the waveform of the decoded audio signal where it is desirable to preserve at least partially (e.g., the temporal part and / or the frequency part). Thus, the possibility of varying the amount of residual signal included in the encoded representation enables a more accurate residual-based reconstruction of the multichannel audio signal. Furthermore, the above-described multichannel audio decoder is between (primarily) parametric encoding and (primarily) residual encoding. Smooth transition In contrast, it does not even require additional signaling, so it should be noted that a very efficient concept is constructed when combined with the multi-channel audio decoder described above. Therefore, the multi-channel encoder described here can take advantage of the benefits made possible by using the multi-channel audio encoder described above.
[0029] In a preferred embodiment, the multichannel audio encoder is configured to vary the bandwidth of the residual signal according to the multichannel audio signal. Thus, it is possible to adjust the residual signal so that it helps to restore the psychoacoustically most important frequency band or frequency range.
[0030] In a preferred embodiment, the multichannel audio encoder is configured to select the frequency bands to which the residual signal should be included in the encoded representation, according to the multichannel audio signal. Thus, the multichannel audio encoder can determine which frequency bands are necessary or most beneficial to include the residual signal (which typically results in at least partial waveform reconstruction). For example, psychoacoustically significant frequency bands can be considered. Additionally, the presence of transient events can also be considered, as the residual signal typically helps improve the rendering of transient phenomena in the audio decoder. Furthermore, the available bitrate can also be taken into account to determine how much residual signal should be included in the encoded representation.
[0031] In a preferred embodiment, the multichannel audio encoder is configured to selectively include residual signals in the encoded representation for frequency bands where the multichannel audio signal is tonal, while excluding the inclusion of residual signals in the encoded representation for frequency bands where the multichannel audio signal is not tonal. This embodiment is based on the consideration that the audio quality obtainable on the audio decoder side can be improved, especially when the tonal frequency bands are reproduced with particularly high quality, preferably when waveform reconstruction is used at least partially. Therefore, selectively including residual signals in the encoded representation for frequency bands where the multichannel audio signal is tonal is beneficial because it results in a good compromise between bitrate and audio quality.
[0032] In a preferred embodiment, the multichannel audio encoder is configured to selectively include residual signals in the encoded representation for the time portion and / or frequency bands resulting from the cancellation of signal components of the multichannel audio signal when forming the downmix signal. It has been found that when there is cancellation of components in the multichannel audio signal, it is difficult or even impossible to properly reconstruct multiple audio signals based on the downmix signal, because the canceled signal components cannot be recovered even by decorrelation or prediction when forming the downmix signal. In such cases, the use of residual signals is an effective way to avoid significant degradation of the reconstructed multichannel audio signal. Thus, this concept helps to improve audio quality while avoiding signaling effort (for example, when combined with the audio decoder described above).
[0033] In a preferred embodiment, the multichannel audio encoder is configured to detect the cancellation of signal components of the multichannel audio signal in the downmix signal, and the multichannel audio decoder is also configured to activate the provision of a residual signal in response to the detection result. Thus, there is an effective way to avoid poor audio quality.
[0034] In a preferred embodiment, the multichannel audio encoder is configured to calculate a residual signal using a linear combination of at least two channel signals of the multichannel audio signal and its dependence on the upmix coefficient used on the multichannel decoder side. Thus, the residual signal is calculated in an efficient manner and is well-suited for the reconstruction of the multichannel audio signal on the multichannel audio decoder side.
[0035] In one embodiment, the multichannel audio encoder is configured to encode upmix coefficients using parameters that describe the inter-channel dependencies of the multichannel audio signal, or to derive upmix coefficients from parameters that describe the inter-channel dependencies of the multichannel audio signal. Thus, the provision of residual signals is parametric. coding It can be efficiently executed based on the parameters used for this purpose as well.
[0036] In a preferred embodiment, the multichannel audio encoder is configured to determine the amount of residual signal included in the encoded representation as a time variable using a psychoacoustic model. Therefore, a relatively high amount of residual signal can be included for portions of the multichannel audio signal with relatively high psychoacoustic relevance (time portions, frequency portions, or time-frequency portions), while a (relatively) smaller amount of residual signal can be included for time portions, frequency portions, or time-frequency portions of the multichannel audio signal with relatively low psychoacoustic relevance. Thus, a good trade-off between bitrate and audio quality can be achieved.
[0037] In a preferred embodiment, the multichannel audio encoder is configured to determine, as a time variable, the amount of residual signal included in the encoded representation according to the currently available bitrate. Thus, the audio quality can be adapted to the available bitrate, making it possible to obtain the best possible audio quality for the currently available bitrate.
[0038] Embodiments of the present invention construct a method for providing at least two output audio signals based on an encoded representation. The method comprises the step of obtaining one of the output audio signals by performing a weighted combination of a downmix signal, an uncorrelated signal, and a residual signal. The weights describing the contribution of the uncorrelated signal in the weighted combination are determined according to the residual signal. This method is based on the same considerations as the audio decoder described above.
[0039] Another embodiment of the present invention constructs a method for providing at least two output audio signals based on an encoded representation. The method is based on an encoded representation of a downmix signal and a plurality of encoded spatial parameters and an encoded representation of a residual signal. , out Power audio signal (At least) one of Steps to obtain include . Residue Difference signal Dependent ,parametric coding and residuals coding Between Blending (or fading) This method is executed. This method is also based on the same considerations as the audio decoder described above.
[0040] Another embodiment of the present invention constructs a method for providing an encoded representation of a multichannel audio signal. The method comprises the steps of: obtaining a downmix signal based on a multichannel audio signal; providing parameters describing the inter-channel dependencies of the multichannel audio signal; and providing a residual signal. The amount of residual signal included in the encoded representation is varied according to the multichannel audio signal. This method is based on the same considerations as the audio encoder described above.
[0041] A further embodiment of the present invention involves constructing a computer program that performs the method described in the present specification. [Brief explanation of the drawing]
[0042] Embodiments of the present invention will be described later with reference to the following drawings. [Figure 1] Figure 1 shows a schematic block diagram of a multi-channel audio encoder according to one embodiment of the present invention. [Figure 2] Figure 2 shows a schematic block diagram of a multi-channel audio decoder according to one embodiment of the present invention. [Figure 3] Figure 3 shows a schematic block diagram of a multichannel audio decoder according to another embodiment of the present invention. [Figure 4] Figure 4 shows a flowchart of a method for providing an encoded representation of a multichannel audio signal according to one embodiment of the present invention. [Figure 5] Figure 5 shows a flowchart of a method for providing at least two output audio signals based on an encoded representation according to one embodiment of the present invention. [Figure 6] Figure 6 shows a flowchart of a method for providing at least two output audio signals based on an encoded representation according to another embodiment of the present invention. [Figure 7] Figure 7 shows a flowchart of a decoder according to one embodiment of the present invention. [Figure 8] Figure 8 shows a schematic representation of the hybrid residual decoder. [Modes for carrying out the invention]
[0043] 1. Multichannel audio encoder shown in Figure 1
[0044] Figure 1 shows a schematic block diagram of a multi-channel audio encoder 100 that provides an encoded representation of a multi-channel signal.
[0045] The multichannel audio encoder 100 is configured to receive a multichannel audio signal 110 and provide an encoded representation 112 of the multichannel audio signal 110 based on it. The multichannel audio encoder 100 also includes a processor (or processing device) 120 configured to receive a multichannel audio signal and obtain a downmix signal 122 based on the multichannel audio signal 110. The processor 120 is further configured to provide parameters 124 that describe the inter-channel dependencies of the multichannel audio signal 110. Furthermore, the processor 120 is configured to provide a residual signal 126. The multichannel audio encoder also includes residual signal processing 130 configured to vary the amount of residual signal contained in the encoded representation 112 according to the multichannel audio signal 110.
[0046] However, it should be noted that a multi-channel audio decoder does not necessarily need to have a separate processor 120 and a separate residual signal processor 130. Rather, it is sufficient if the multi-channel audio encoder is configured in some way to perform the functions of the processor 120 and the residual signal processor 130.
[0047] Regarding the function of the multi-channel audio encoder 100, the channel signals of the multi-channel audio signal 110 are Typically It is encoded using multi-channel coding. This was pointed out, and here Encoded representation 112 is (encoded form of ) The downmix signal 122, the parameter 124 describing the dependency between channels (or channel signals) of the multichannel audio signal 110, and the residual signal 126 Typically To prepare, the downmix signal 122 may be based on, for example, the combination (e.g., linear combination) of channel signals of a multichannel audio signal. aHowever, the downmix signal 122 is provided based on the multiple channel signals of the multichannel audio signal. So that a ru. however , or two or more downmixed signals but , multichannel audio signal 110 many of (typically) The number of downmix signals is greater than many ) Channel signal Related to Moa Parameter 124 can describe the dependencies (e.g., correlation, covariance, level relationships, etc.) between channels (or channel signals) of the multi-channel audio signal 110. Therefore, parameter 124 is intended to allow the audio decoder to derive a restored version of the channel signals of the multi-channel audio signal 110 based on the downmix signal 122. to accomplish This purpose for , parameter 124 is Describe the desired characteristics (e.g., individual or relative characteristics) of the channel signals in a multi-channel audio signal. This allows an audio encoder using parametric decoding to reconstruct the channel signal based on one or more downmix signals 122. do.
[0048] In addition, the multi-channel audio decoder 100 provides a residual signal 126 that typically represents signal components that cannot be restored by the audio decoder (e.g., an audio decoder following specific processing rules) based on the downmix signal 122 and parameters 124, as predicted or estimated by the multi-channel audio encoder. Therefore, the residual signal 126 can typically be considered an improved signal that enables waveform restoration, or at least partial waveform restoration, on the audio decoder side.
[0049] However, the multichannel audio encoder 100 is configured to vary the amount of residual signal included in the encoded representation 112 according to the multichannel audio signal 110. In other words, the multichannel audio encoder can determine, for example, the intensity (or energy) of the residual signal 126 included in the encoded representation 112. In addition, or alternatively, the multichannel audio encoder 100 can determine for which frequency bands and / or how many frequency bands the residual signal is included in the encoded representation 112. By varying the "amount" of residual signal 126 included in the encoded representation 112 according to the multichannel audio signal (and / or the available bitrate), the multichannel audio encoder 100 can flexibly determine with what accuracy the channel signals of the multichannel audio signal 110 can be reconstructed on the audio decoder side based on the encoded representation 112. Thus, the accuracy with which the channel signals of the multichannel audio signal 110 can be reconstructed can be adapted to the psychoacoustic relationships of different signal parts of the channel signals of the multichannel audio signal 110 (e.g., time part, frequency part and / or time / frequency part). Therefore, signal portions with high psychoacoustic relevance (e.g., tonal signal portions or signal portions with transient events) can be encoded with particularly high resolution by including a "large" amount of residual signal 126 in the encoded representation. For example, for signal portions with high psychoacoustic relevance, it can be achieved that residual signals with relatively high energy are included in the encoded representation 112. Furthermore, if the downmix signal 122 contains "low quality" signals, for example, when the channel signals of a multi-channel audio signal 112 are coupled to the downmix signal 122, and there is substantial cancellation of signal components, it can be achieved that high-energy residual signals are included in the encoded representation 112.In other words, the multi-channel audio decoder 100 can selectively embed a "larger amount" of residual signal (e.g., a residual signal with relatively high energy) into the encoded representation of the signal portion of the multi-channel audio signal 110, where providing a relatively large amount of residual signal results in a significant improvement in the restored channel signal (restored by the audio decoder).
[0050] Therefore, the change in the amount of residual signal included in the encoded representation according to the multichannel audio signal 110 makes it possible to adapt the encoded representation 112 of the multichannel audio signal 110 (for example, the residual signal 126 included in the encoded representation in encoded form) so that a good trade-off can be achieved between bitrate efficiency and the audio quality of the restored multichannel audio signal (restored on the audio decoder side).
[0051] It should be noted that the multi-channel audio encoder 100 can be improved as an option in many different ways. For example, the multi-channel audio encoder can be configured to vary the bandwidth of the residual signal 126 (included in the encoded representation) according to the multi-channel audio signal 110. Thus, the amount of residual signal included in the encoded representation 112 can be adapted to the perceptually most important frequency band.
[0052] Optionally, the multichannel audio decoder can be configured to select the frequency bands in which the residual signal 126 is contained in the encoded representation 112 according to the multichannel audio signal 110. Thus, the encoded representation 120 (more precisely, the amount of residual signal contained in the encoded representation 112) can be adapted to the multichannel audio signal, for example, to the perceptually most important frequency bands of the multichannel audio signal 110.
[0053] Optionally, the multichannel audio encoder can be configured to include residual signals 126 in the encoded representation for frequency bands where the multichannel audio signal is tonal. In addition, the multichannel audio encoder can be configured not to include residual signals 126 in the encoded representation 112 for frequency bands where the multichannel audio signal is not tonal (unless any other specific conditions are met that would result in the inclusion of residual signals in the encoded representation for a particular frequency band). Thus, residual signals can be selectively included in the encoded representation for perceptually important tonal frequency bands.
[0054] Optionally, the multichannel audio encoder 100 can be configured to selectively include residual signals in the encoded representation for time portions and / or frequency bands where the formation of a downmix signal results in the cancellation of signal components of the multichannel audio signal. For example, the multichannel audio encoder can be configured to detect the cancellation of signal components of the multichannel audio signal 110 in the downmix signal 122 and activate the provision of residual signals 126 (e.g., inclusion of residual signals 126 in the encoded representation 112) according to the result of the detection. Thus, if the downmix of the channel signals of the multichannel audio signal 110 to the downmix signal 122 (or any other normal linear combination) results in the cancellation of signal components of the multichannel audio signal 112 (which may be caused, for example, by signal components of different channel signals that are 180 degrees phase-shifted), the encoded representation 112 includes residual signals 126 that help overcome the detrimental effects of this cancellation when the audio decoder restores the multichannel audio signal 110. For example, the residual signal 126 can be selectively included in the encoded representation 112 for frequency bands where such cancellations occur.
[0055] Optionally, the multichannel audio encoder can be configured to calculate the residual signal using a linear combination of at least two channel signals of the multichannel audio signal, according to the upmix coefficient used on the multichannel audio decoder side. Such calculation of the residual signal is efficient and allows for easy reconstruction of the channel signals on the audio decoder side.
[0056] Optionally, the multichannel audio encoder can be configured to encode the upmix coefficient using parameter 124, which describes the inter-channel dependency of the multichannel audio signal, or to derive the upmix coefficient from a parameter that describes the inter-channel dependency of the multichannel audio signal. 、( example Bachi Channel level difference parameter, channel correlation parameter, etc. of ) Parameter 124 is parametric coding (Encoding or Decoding) and Residual signal assist coding It can be used for both encoding and decoding. In this way , use of residual signal 126 by , additional Signaling Overhead but Motarasa Being No. Rather, it's parametric in any case. coding The parameter 124 used for (encoding / decoding) is the residual. coding It is also reused for (encoding / decoding). Therefore This allows for the achievement of high coding efficiency.
[0057] Optionally, a multi-channel audio decoder can be configured to use a psychoacoustic model to determine the amount of residual signal included in the encoded representation as a time variable. Thus, the encoding accuracy can be adapted to the psychoacoustic characteristics of the signal, which usually results in good bitrate efficiency.
[0058] However, it should be noted that a multi-channel audio encoder can be optionally supplemented by any of the features or functions described in this specification (both the specification and the claims). Furthermore, a multi-channel audio encoder can be adapted in parallel with the audio decoder described in this specification in order to work with the audio decoder.
[0059] 2. Multichannel audio decoder shown in Figure 2
[0060] Figure 2 shows a schematic block diagram of a multi-channel audio decoder 200 according to one embodiment of the present invention.
[0061] The multichannel audio decoder 200 is configured to receive an encoded representation 210 and provide at least two output audio signals 212, 214 based thereon. The multichannel audio decoder 200 may include a weighted coupler 220 configured to perform a weighted coupling of a downmix signal 222, an uncorrelated signal 224, and a residual signal 226 to obtain (at least) one output signal, for example, a first output audio signal 212. It should be noted that the downmix signal 212, the uncorrelated signal 224, and the residual signal 226 can be derived, for example, from the encoded representation 210, and the encoded representation 210 may be accompanied by an encoded representation of the downmix signal 220 and an encoded representation of the residual signal 226. Furthermore, the uncorrelated signal 224 can be derived, for example, from the downmix signal 222, or can be derived using additional information contained in the encoded representation 210. However, the uncorrelated signal can also be provided without any dedicated information from the encoded representation 210.
[0062] The multichannel audio decoder 200 is also configured to determine weights describing the contribution of the uncorrelated signal 224 in the weighted coupling, according to the residual signal 226. For example, the multichannel audio decoder 200 may include a weight determination unit 230 configured to determine weights 232 describing the contribution of the uncorrelated signal 224 in the weighted coupling (e.g., the contribution of the uncorrelated signal 224 to the first output audio signal 212) based on the residual signal 226.
[0063] Regarding the functionality of the multi-channel audio decoder 200, it should be noted that the contribution of the uncorrelated signal 224 to the weighted coupling, and consequently to the first output audio signal 212, is adjusted in a flexible manner (e.g., in a time-variable and frequency-dependent manner) according to the residual signal 226 without additional signaling overhead. Thus, the amount of uncorrelated signal 224 included in the first output audio signal 212 is adapted according to the amount of residual signal 226 included in the first output audio signal 212 so that good quality of the first output audio signal 212 is achieved. Therefore, it is possible to obtain appropriate weighting of the uncorrelated signal 224 without additional signaling overhead under any circumstances. Thus, good quality of the decoded output audio signal 212 can be achieved at a moderate bitrate using the multi-channel audio decoder 200. The accuracy of the reconstruction can be flexibly adjusted by the audio encoder, which can determine the amount of residual signal 226 contained in the encoded representation 210 (e.g., how large the energy of the residual signal 226 contained in the encoded representation 210 is, or to what frequency band the residual signal 226 contained in the encoded representation 210 relates to), and the multi-channel audio decoder 200 can react accordingly and adjust the weighting of the uncorrelated signal 224 to fit the amount of residual signal 226 contained in the encoded representation 210. Consequently, when there is a large amount of residual signal 226 contained in the encoded representation 210 (e.g., for a particular frequency band or for a particular time portion), the weighted coupling 220 can primarily (or exclusively) consider the residual signal 226, while little (or no) weight is given to the uncorrelated signal 224. In contrast, when the encoded representation 210 contains only a smaller amount of residual signal 226, the weighted concatenation 220 can consider the downmix signal 222, as well as primarily (or exclusively) the uncorrelated signal 224, but the residual signal 226 is given only a relatively small amount of weight (or no weight at all).Therefore, the multi-channel audio decoder 200 can flexibly work with a suitable multi-channel audio encoder and adjust the weighted coupling 220 to achieve the best audio quality under any circumstances (regardless of whether a smaller or larger amount of residual signal 226 is included in the encoded representation 210).
[0064] It should be noted that the second output audio signal 214 can be generated in a similar manner. However, if there are different quality requirements for the second output audio signal, for example, it is not always necessary to apply the same mechanism to the second output audio signal 214.
[0065] In an optional improvement, the multi-channel audio decoder can be configured to determine weights 232 that describe the contribution of the uncorrelated signal 224 to the weighted coupling, according to the uncorrelated signal 224. In other words, weights 232 can be dependent on both the residual signal 226 and the uncorrelated signal 224. Thus, weights 232 can even be better fitted to the current decoded audio signal without additional signaling overhead.
[0066] As an improvement to other options, the multichannel audio decoder can be configured to obtain upmix parameters based on the encoded representation 212 and determine weights 232 that describe the contribution of the uncorrelated signals in the weighted combination according to the upmix parameters. Thus, the weights 232 can be additionally dependent on the upmix parameters to achieve a better fit of the weights 232.
[0067] another of any As an improvement, the multi-channel audio decoder , heavy The weights that describe the contribution of the uncorrelated signal in the matched coupling are The weights of the uncorrelated signal decrease as the energy of the residual signal increases. It can be configured to make a decision. Therefore, (In addition to the downmix signal 222) Mainly uncorrelated signal 22 4 Decryption based on, (In addition to the downmix signal 222) Mainly residual signal 22 6 Between the decoding based on Blend or fade It can be executed.
[0068] another of any As an improvement, the multi-channel audio decoder 200 is configured such that the energy of the residual signal 226 is zero. to ( Included in encoded representation 210 Meru (It is possible, or it can be derived from it.) Uncorrelated signal upmix parameters The energy of the residual signal 226, weighted by the residual signal weighting coefficient (or residual signal upmix parameter), is such that the maximum weight determined by is related to the uncorrelated signal 224. No Energy of the uncorrelated signal 224 weighted by the correlated signal upmix parameter That's all. In some cases such The weights 232 can be determined such that the zero weights correspond to the uncorrelated signal 204. Thus, there is a perfect relationship between decoding based on the uncorrelated signal 224 and decoding based on the residual signal 226. blend do (or Fading It is possible to do this if the residual signal 226 is judged to be sufficiently strong (for example, if the energy of the weighted residual signal is equal to the energy of the weighted uncorrelated signal 224). That's all. time )、 Weighted joins When improving the downmix signal 222 Without considering the uncorrelated signal 224, Completely Residual signal 226 by It can be kept. In this case, The multi-channel audio decoder 200 can perform particularly good (at least partial) waveform reconstruction. Uncorrelated signal 224 of consideration Typically, this is done This is particularly good waveform restoration. but Obstruction Rare In contrast, residual signal 226 of use Typically, this is done teeth especially Good waveform restoration but Possible Na ru Therefore .
[0069] In other optional improvements, the multichannel audio decoder 200 can be configured to calculate weighted energy values for an uncorrelated signal weighted according to one or more uncorrelated signal upmix parameters, and to calculate weighted energy values for a residual signal weighted using one or more residual signal upmix parameters. In this case, the multichannel audio decoder can be configured to determine elements according to the weighted energy values for the uncorrelated signal and the weighted energy values for the residual signal, and to obtain weights based on these elements that describe the contribution of the uncorrelated signal 224 to one output audio signal (e.g., a first output audio signal 212). Thus, the weight determination 230 can provide particularly well-fitted weight values 232.
[0070] In an optional improvement, the multichannel audio decoder 200 (or its weighting determiner 230) may be configured to multiply its elements by an uncorrelated signal upmix parameter (which may be included in or derived from the encoded representation 210) in order to obtain a weight (or weighting value) 232 that describes the contribution of the uncorrelated signal 224 to one output audio signal (e.g., a first output audio signal 212).
[0071] In an optional improvement, the multichannel audio decoder (or its weighting determiner 230) may be configured to calculate the energy of the uncorrelated signal 224 weighted with uncorrelated signal upmix parameters (which may be included in or derived from the encoded representation 210) across multiple upmix channels and time slots in order to obtain the weighted energy value of the uncorrelated signal 224.
[0072] As a further option for improvement, the multi-channel audio decoder 200 can be configured to calculate the energy of the residual signal 224 weighted using residual signal upmix parameters (which may be included in or derived from the encoded representation 210) across multiple upmix channels and time slots in order to obtain weighted energy values of the residual signal.
[0073] As an improvement to other options, the multi-channel audio decoder 200 (or its weighting determiner 232) can be configured to calculate the aforementioned elements according to the difference between the weighted energy values of the uncorrelated signal and the weighted energy values of the residual signal. Such calculations have been found to be an efficient solution for determining the weighting values 232.
[0074] As an optional improvement, the multichannel audio decoder can be configured to calculate the elements according to the ratio of the difference between the weighted energy value of the uncorrelated signal 224 and the weighted energy value of the residual signal 226, and the weighted energy value of the uncorrelated signal 224. Such calculations for the elements are between the improvement of the downmix signal 222 mainly based on the uncorrelated signal and the improvement of the downmix signal 222 mainly based on the residual signal. Smooth transition It has been shown to yield positive results in this regard.
[0075] As an optional improvement, the multi-channel audio decoder 200 can be configured to determine weights describing the contribution of the uncorrelated signal to two or more output audio signals, such as a first output audio signal 212 and a second output audio signal 214. In this case, the multi-channel audio decoder can be configured to determine the contribution of the uncorrelated signal 224 to the first output audio signal 212 based on the weighted energy value of the uncorrelated signal 224 and the uncorrelated signal upmix parameter of the first channel. Furthermore, the multi-channel audio decoder can be configured to determine the contribution of the uncorrelated signal 224 to the second output audio signal 214 based on the weighted energy value of the uncorrelated signal 224 and the uncorrelated signal upmix parameter of the second channel. In other words, different uncorrelated signal upmix parameters can be used to provide the first output audio signal 212 and the second output audio signal 214. However, the same weighted energy value of the uncorrelated signal can be used to determine the contribution of the uncorrelated signal to the first output audio signal 212 and the contribution of the uncorrelated signal to the second output audio signal 214. Therefore, effective adjustments can be made that can be considered by different uncorrelated signal upmix parameters, regardless of the different characteristics of the two output audio signals 212 and 214.
[0076] As an optional improvement, the multi-channel audio decoder 200 can be configured to disable the contribution of the uncorrelated signal 224 to the weighted coupling if the residual energy (e.g., the energy of the residual signal 226 or the energy of the weighted version of the residual signal 226) exceeds the uncorrelated energy (e.g., the energy of the uncorrelated signal 224 or the energy of the weighted version of the uncorrelated signal 224).
[0077] As a further option for improvement, the audio decoder can be configured to determine, band by band, weights 232 describing the contribution of the uncorrelated signal 224 in the weighted coupling, according to the band by band determination of the weighted energy values of the residual signal. Thus, fine-tuning of the multi-channel audio decoder 200 to the signal being decoded can be performed.
[0078] In other optional improvements, the audio decoder can be configured to determine weights for each frame of the output audio signals 212,214 that describe the contribution of the uncorrelated signals in the weighted combination. Thus, good temporal resolution can be achieved.
[0079] In further refinement of the options, the determination of the weight value 232 can be carried out by several formulas provided below.
[0080] Furthermore, it should be noted that the multi-channel audio decoder 200 can be supplemented with any of the features or functions described in this specification in other embodiments as well.
[0081] 3. Multichannel audio decoder shown in Figure 3
[0082] Figure 3 shows a schematic block diagram of a multichannel audio decoder 300 according to one embodiment of the present invention. The multichannel audio decoder 300 is configured to receive an encoded representation 310 and provide two or more output audio signals 312, 314 based thereon. The encoded representation 310 may include, for example, an encoded representation of a downmix signal, an encoded representation of one or more spatial parameters, and an encoded representation of a residual signal. The multichannel audio decoder 300 is configured to obtain (at least) one output audio signal, for example, a first output audio signal 312 and / or a second output audio signal 314, based on the encoded representation of the downmix signal, a plurality of encoded spatial parameters, and an encoded representation of a residual signal.
[0083] In particular, the multi-channel audio decoder 300 , (encoded encoded representation 310 Include (Rare) Residual signal to Dependent ,parametric coding and residuals coding Between blend It is configured to provide output audio signals 312, 314. In other words, the multi-channel audio decoder 300 provides output audio signals 312, 314. But Based on the UNMIX signal, and Desired relationship between output audio signals 312 and 314 Person in charge ( For example, the desired inter-channel level difference or desired inter-channel correlation of output audio signals 312 and 314. Spatial parameters that describe The decoding mode is performed using and the output audio signals 312, 314 But Based on the hummix signal Using residual signals Between the decryption mode being restored blend It is possible. In this way The intensity (e.g., energy) of the residual signal included in the coded representation 310 is , revenge Numbering generally (or solely)( In addition to the downmix signal ) Spatial parameters Whether it is based on or the decryption is generally (or solely)( (In addition to the downmix signal) Residual signal Determine whether it is based on this, or whether an intermediate state is taken where both spatial parameters and residual signals affect the improvement of the downmix signal. Then, the output audio signals 312 and 314 are derived from the downmix signal. It is possible.
[0084] moreover, Depending on the residual signal, Multi-channel audio decoder 300 by , (typically) When providing output audio signals 312, 314 No For the correlated signal A relatively large weight (given) Parametric coding and (typically teeth , nothing Correlated signals In contrast, relatively small weight (given) Residual coding Between blend This allows for decoding that is well-suited to current audio content without high signaling overhead. but Possible Na ru.
[0085] Furthermore, it should be noted that the multi-channel audio decoder 300 is based on similar considerations to the multi-channel audio decoder 200, and the optional improvements described above for the multi-channel audio decoder 200 can also be applied to the multi-channel audio decoder 300.
[0086] 4. A method for providing an encoded representation of a multichannel audio signal as shown in Figure 4.
[0087] Figure 4 shows a flowchart of method 400, which provides an encoded representation of a multichannel audio signal.
[0088] Method 400 comprises step 410 of obtaining a downmix signal based on a multichannel audio signal. Method 400 also comprises step 420 of providing parameters that describe the inter-channel dependencies of the multichannel audio signal. For example, inter-channel level difference parameters and / or inter-channel correlation parameters (or covariance parameters) that describe the inter-channel dependencies of the multichannel audio signal can be provided. Method 400 also comprises step 430 of providing a residual signal. Furthermore, Method 440 comprises step 440 of varying the amount of residual signal included in the encoded representation according to the multichannel audio signal.
[0089] It should be noted that Method 400 is based on the same considerations as the audio encoder 100 shown in Figure 1. Furthermore, Method 400 can be supplemented by any of the features and functions described in this specification with respect to the apparatus of the present invention.
[0090] 5. A method for providing at least two output audio signals based on the encoded representation shown in Figure 5.
[0091] Figure 5 shows a flowchart of method 500 for providing at least two output audio signals based on an encoded representation. Method 500 comprises a step 510 for determining weights that describe the contribution of the uncorrelated signal in the weighted combination according to the residual signal. Method 500 also comprises a step 520 for obtaining one of the output audio signals by performing a weighted combination of the downmix signal, the uncorrelated signal and the residual signal.
[0092] It should be noted that Method 500 can be supplemented by any of the features and functions described in the present specification with respect to the apparatus of the present invention.
[0093] 6. A method for providing at least two output audio signals based on the encoded representation shown in Figure 6.
[0094] Figure 6 shows that at least two output audio signals are provided based on the encoded representation. for A flowchart of Method 600 is shown. Method 600 modifies the output audio signal based on the encoded representation of the downmix signal and the encoded representation of the residual signal and multiple encoded spatial parameters. Our Step 610 to get one include Output audio signal Our Step 610 to obtain one is the residual signal Dependent ,parametric coding and residuals coding Between blend Step 620 to execute include .
[0095] It should be noted that Method 600 can be supplemented by any of the features and functions described in the present specification with respect to the apparatus of the present invention.
[0096] 7. Further Embodiments
[0097] Some general considerations and some further embodiments are described below.
[0098] 7.1 General Considerations
[0099] The embodiment of the present invention uses a decoder (e.g., a multi-channel audio decoder) instead of a fixed residual bandwidth. but , in each frame Regarding (Or, generally, in at least multiple frequency ranges) Regarding and / or multiple time segments Regarding ) Ba By measuring the energy for each unit 、 It is based on the idea of detecting the amount of transmitted residual signal. Depending on the transmitted spatial parameters, the uncorrelated output is added where the residual energy is "lost" to obtain the output energy and the required (or desired) amount of uncorrelatedness. by Similar to bandpass-style residual signals, the residual bandwidth is variable. but Possible Na For example, the residuals relative to the tonal band coding Use only too It is possible. (Also known as residual coding) Waveform storage coding Similar to the parametric coding To enable the use of a simplified downmix, a residual signal for the simplified downmix is defined in this specification.
[0100] 7.2 Calculation of residual signals for a simplified downmix
[0101] The following sections discuss the calculation of residual signals and some considerations regarding the structure of channel signals in multi-channel audio signals.
[0102] In Unified Speech and Audio Coding (USAC), there is no defined residual signal when so-called "simplified downmix" is used. Therefore, no partial waveform preservation coding is possible. However, the following describes a method for calculating the residual signal for so-called "simplified downmix".
[0103] Parametric upmix coefficient u d1 ,u d2 While the parameter weights are calculated for each parameter band, the "simplified downmix" weights d1 and d2 are calculated for each scale coefficient band. Therefore, the coefficient w used to calculate the residual signal is calculated separately. r1 ,w r2 This cannot be calculated directly from the spatial parameters (as it is the case for classical MPEG surround), but it may require being determined for each scale factor band from the downmix and upmix coefficients.
[0104] Here, if L and R are input channels and D is the downmix channel, the residual signal res must satisfy the following characteristics. TIFF0007862341000001.tif19145
[0105] This is achieved by calculating the residuals as follows: TIFF0007862341000002.tif8145 Here, we use the following downmix weights. TIFF0007862341000003.tif26145
[0106] The residual upmix coefficient u used by the decoder r,1 ,u r,2 The method is preferably selected in a way that ensures robust decoding. Since the simplified downmix has asymmetric properties (in contrast to fixed-weighted MPEG surround), a spatial parameter-dependent upmix is applied, for example, using the following upmix coefficients. TIFF0007862341000004.tif15145
[0107] The other option is to define a residual upmix coefficient that is orthogonal to the upmix coefficient of the downmix signal as follows. TIFF0007862341000005.tif14146
[0108] In other words, the audio decoder can obtain the downmix signal D using a linear combination of the left channel signal L (the first channel signal) and the right channel signal R (the second channel signal). Similarly, the residual signal res is obtained using a linear combination of the left channel signal L and the right channel signal R (or, generally, the first channel signal and the second channel signal of a multi-channel audio signal).
[0109] For example, in equations (5) and (6), when the simplified downmix weights d1, d2, the parametric upmix coefficients u d,1 ,u d,2 and the residual upmix coefficients u r,1 ,u r,2 are determined, it can be seen that the downmix weights w r,1 ,w r,2 for obtaining the residual signal res can be obtained. Furthermore, it can be seen that u r,1 ,u r,2 can be derived from u d,1 ,u d,2 using equations (7) and (8) or equation (9). The simplified downmix weights d1, d2 can be obtained in the usual way, similar to the parametric upmix coefficients u d,1 ,u d,2 .
[0110] 7.3 Encoding Process
[0111] The following details some of the encoding process. Encoding can be performed, for example, by a multi-channel audio encoder 100, or by any other suitable means or computer program.
[0112] Preferably, the amount of transmitted residual is determined by a psychoacoustic model of an encoder (e.g., a multi-channel audio encoder), depending on the audio signal (e.g., the channel signals of a multi-channel audio signal 110) and the available bitrate. The transmitted residual signal can be used, for example, for partial waveform preservation or to avoid signal cancellation caused by the downmixing method used (e.g., the downmixing method described by equation (1) above).
[0113] 7.3.1 Partial waveform storage
[0114] The following describes how partial waveform preservation can be achieved. For example, the calculated residual (e.g., residual res by equation (4)) is transmitted in full band or band-limited to provide partial waveform preservation within the residual bandwidth. The residual portion, which is perceived as perceptually irrelevant by the psychoacoustic model, can be quantized to zero (e.g., based on the residual signal 126 when providing the encoded representation 112). This includes, but is not limited to, reducing the transmitted residual bandwidth at runtime (which can be thought of as changing the amount of residual signal contained in the encoded representation). The system can also allow for bandpass-style erasure of the residual signal portion, since the lost signal energy is restored by a decoder (e.g., a multi-channel audio decoder 200 or a multi-channel audio decoder 300). Thus, while background noise can be parametrically encoded to reduce the residual bitrate, residual coding can, for example, be applied only to the tonal components of the signal while preserving their phase relationships. In other words, the residual signal 126 can be included in the encoded representation 112 only for frequency bands and / or time portions where the multichannel audio signal 110 (or at least one channel signal of the multichannel audio signal 110) is found to be tonal (e.g., by residual signal processing 130). In contrast, the residual signal 126 can be omitted from the encoded representation 112 for frequency bands or time portions where the multichannel audio signal 110 (or at least one or more channel signals of the multichannel audio signal 110) is identified as noise-like. Thus, the amount of residual signal included in the encoded representation varies according to the multichannel audio signal.
[0115] 7.3.2 Preventing signal cancellation in downmixing
[0116] The following describes how signal cancellation in downmixing can be prevented (or compensated for).
[0117] For low-bitrate applications, parametric coding (which relies primarily or exclusively on a parameter 124 describing the inter-channel dependencies of a multi-channel audio signal) is applied instead of waveform-preserving coding (which relies primarily on a residual signal 126 in addition to, for example, a downmix signal 122). Here, the residual signal 126 is used only to compensate for signal cancellation in the downmix 122 in order to minimize the bit usage of the residual. Unless signal cancellation is detected in the downmix 122, the system operates in parametric mode (on the audio decoder side) using a decorrelator. For example, when signal cancellation occurs for a fading tonal signal, the residual signal 126 is transmitted for the faulty signal portion (e.g., frequency band and / or time portion). Thus, the signal energy can be recovered by the decoder.
[0118] 7.4 Decryption Process
[0119] 7.4.1 Overview
[0120] In the decoder (e.g., multi-channel audio decoder 200 or multi-channel audio decoder 300), the transmitted downmix and residual signals (e.g., downmix signal 222 or residual signal 226) are decoded by the core decoder and fed to the MPEG surround decoder along with the decoded MPEG surround payload. The residual upmix coefficient for a classical MPS downmix is invariant, and the residual upmix coefficient for a simplified downmix is defined by equations (7), (8), and / or (9). In addition, the output of the uncorrelatedizer and its weighting coefficients are calculated with respect to parametric decoding. The residual signal and the output of the uncorrelatedizer are weighted, and both are mixed into the output signal. Thus, the weighting factor is determined by measuring the energy of the residual and uncorrelatedizer signals.
[0121] In other words, the residual upmix element (or coefficient) can be determined by measuring the energy of the residual and uncorrelated signals.
[0122] For example, the downmix signal 222 is provided based on the encoded representation 210, and the uncorrelated signal 224 is generated based on parameters derived from the downmix signal 222 or contained in the encoded representation 210 (or otherwise). The residual upmix coefficient is a parametric upmix coefficient u, generated by the decoder according to, for example, equations (7) and (8). d,1 ,u d,2 It can be derived from the parametric upmix coefficient u d,1 ,u d,2 These can be obtained based on the encoded representation 210, for example, by directly deriving them from spatial data contained in the encoded representation 210 (e.g., from inter-channel correlation coefficients and inter-channel level difference coefficients, or from inter-object correlation coefficients and inter-object level differences).
[0123] Upmix coefficients for the uncorrelatedizer output (or output) can be obtained for conventional MPEG surround decoding. However, weighting coefficients for the weighting of the uncorrelatedizer output (or output(s)) can be determined based on the energy of the residual signal (and possibly also based on the energy of the uncorrelatedizer signal or signal), such that the weights describing the contribution of the uncorrelated signal in the weighted combination are determined according to the residual signal.
[0124] 7.4.2 Exemplary Embodiments
[0125] Exemplary embodiments are described below with reference to Figure 7. However, it should be noted that the concepts described in this specification are applicable to the multi-channel audio decoder 200 or 300 shown in Figures 2 and 3.
[0126] Figure 7 shows a schematic block diagram (or flow diagram) of a decoder (e.g., a multi-channel audio decoder). The decoder in Figure 7 is represented as 700. Decoder 700 is configured to receive a bitstream 710 and provide a first output channel signal 712 and a second output channel signal 714 based thereon. Decoder 700 includes a core decoder 720 configured to receive the bitstream 710 and provide a downmix signal 722, a residual signal 724, and spatial data 726 based thereon. For example, the core decoder 720 can provide a time-domain or transformation-domain representation (e.g., frequency-domain representation, MDCT-domain representation, QMF-domain representation) of the downmix signal represented by bitstream 710 as the downmix signal. Similarly, the core decoder 720 can provide a time-domain or transformation-domain representation of the residual signal 724 represented by bitstream 710. Furthermore, the core decoder 720 can provide one or more spatial parameters 726, such as one or more inter-channel correlation parameters, inter-channel level difference parameters, etc.
[0127] The decoder 700 also includes a decorrelator 730 configured to provide a decorrelated signal 732 based on the downmix signal 722. Any well-known decorrelation concept can be used by the decorrelator 730. Furthermore, the decoder 700 also receives spatial data 726 and upmix parameters (e.g., upmix parameter u dmx,1 ,u dmx,2 ,u dec,1 ,u dec,2 The decoder 700 further comprises an upmix coefficient calculator 740 configured to provide the upmix parameters 742 (also referred to as upmix coefficients) provided by the upmix coefficient calculator 740 based on spatial data 726. For example, the upmixer 750 obtains two upmixed versions 752, 754 of the downmix signal 722 by taking the upmix coefficients (e.g., u dmx,1 ,u dmx,2 The downmix signal 722 can be scaled using ). Furthermore, the upmixer 750 is configured to apply one or more upmix parameters (e.g., two upmix parameters) to the uncorrelated signal 732 provided by the uncorrelatedizer 730 in order to obtain a first upmixed (scaled) version 756 and a second upmixed (scaled) version 758 of the uncorrelated signal 732. Furthermore, the upmixer 750 is configured to apply one or more upmix coefficients (e.g., two upmix coefficients) to the residual signal 724 in order to obtain a first upmixed (scaled) version 760 and a second upmixed (scaled) version 762 of the residual signal 724.
[0128] The decoder 700 also includes a weight calculator 770 configured to measure the energies of upmixed (scaled) versions 756,758 of the uncorrelated signal 752 and the energies of upmixed (scaled) versions 760,762 of the residual signal 724. Furthermore, the weight calculator 770 is configured to provide one or more weight values 772 to a weighter 780. The weighter 780 is configured to use one or more weight values 772 provided by the weight calculator 770 to obtain a first upmixed (scaled), weighted version 782 of the uncorrelated signal 732, a second upmixed (scaled), weighted version 784 of the uncorrelated signal 732, a first upmixed (scaled), weighted version 786 of the residual signal 724, and a second upmixed (scaled), weighted version 788 of the residual signal 724. The decoder also includes a first adder 790 configured to sum a first upmixed (scaled) version 752 of the downmix signal 720, a first upmixed (scaled) and weighted version 782 of the uncorrelated signal 732, and a first upmixed (scaled) and weighted version 786 of the residual signal 724, in order to obtain a first output channel signal 712. Furthermore, the decoder includes a second adder 792 configured to sum a second upmixed version 754 of the downmix signal 720, a second upmixed (scaled) and weighted version 784 of the uncorrelated signal 732, and a second upmixed (scaled) and weighted version 788 of the residual signal 724, in order to obtain a second output channel signal 714.
[0129] However, it should be noted that the weighter 780 does not need to weight all signals 756, 758, 760, and 762. For example, in some embodiments, it may suffice to weight only signals 756 and 758, leaving signals 760 and 762 unaffected (practically, so that signals 760 and 762 are directly applied to adders 790 and 792). Alternatively, however, the weighting of residual signals 760 and 762 can be varied over time. For example, the residual signals can be faded in or faded out. For example, the weighting (or weighting coefficient) of the uncorrelated signals can be smoothed over time, and the residual signals can be faded in or faded out accordingly.
[0130] Furthermore, it should be noted that the weighting performed by the weighter 780 and the upmixing applied by the upmixer 750 can be performed as a combined operation, and the weight calculation can be performed directly using the uncorrelated signal 732 and the residual signal 724.
[0131] Some further details regarding the functionality of Decoder 700 are described below.
[0132] The combined residual and parametric coding modes can be signaled, for example, in a quasi-backward compatible manner, by signaling the residual bandwidth of one parameter band in the bitstream. Thus, the legacy decoder still passes through and decodes the bitstream by switching to parametric decoding on the first parameter band. The legacy bitstream with residual bandwidth does not contain residual energy on the first parameter band and becomes parametrically decoded in the proposed novel decoder.
[0133] However, within a 3D audio codec system, combined residual and parametric coding can be used in combination with other core decoder tools such as quad-channel elements, allowing the decoder to explicitly detect legacy bitstreams and enable their decoding in the normal band-limited residual coding mode. The actual residual bandwidth is determined by the decoder at runtime and is therefore preferably not explicitly signaled. The calculation of the upmix coefficients is set to parametric mode instead of residual coding mode. The energy Edec of the weighted uncorrelatedizer output and the energy of the weighted residual signal Eres are calculated for each upmix channel ch for each frame and the hybrid band hb across all time slots ts, as follows: TIFF0007862341000006.tif22147
[0134] TIFF0007862341000007.tif48168
[0135] The residual signals (e.g., upmixed residual signal 760 or upmixed residual signal 762) are added to the output channels (e.g., output channels 712, 714) with a weight of 1. The uncorrelated signals (e.g., upmixed uncorrelated signal 756 or upmixed uncorrelated signal 758) can be weighted by an element r (e.g., by the weighter 780) calculated as follows: TIFF0007862341000008.tif21168 Here, E dec (hb) is the uncorrelated signal x for the frequency band hb. dec This represents the weighted energy value of E res (hb) is the residual signal x for the frequency band hb. res This represents the weighted energy value.
[0136] If the residual (for example, residual signal 724) is not transmitted, for example, E resIf = 0, then r (a coefficient that can be applied by the weighter 780 and can be considered as the weighted value 772) becomes 1, which is purely parametric decoding. If the residual energy (e.g., the energy of the upmixed residual signal 760 and / or upmixed residual signal 762) exceeds the uncorrelatedizer energy (e.g., the energy of the upmixed uncorrelated signal 756 or upmixed uncorrelated signal 758), then, for example, E res >E dec In this case, the coefficient r can be set to zero, thus disabling the decorrelator and enabling partial waveform preservation decoding (which can be considered residual coding). In the upmixing process, both the weighted decorrelator output (e.g., signals 782,784) and the residual signal (e.g., signals 786,788 or signals 760,762) are added to the output channel (e.g., signals 712,714).
[0137] In conclusion, this becomes a matrix-style upmix rule. TIFF0007862341000009.tif18154 Here, ch1 represents one or more time-domain samples or transformed-domain samples of the first output audio signal, ch2 represents one or more time-domain samples or transformed-domain samples of the second output audio signal, x dmx x represents one or more time-domain samples or transformed-domain samples of the downmix signal, dec x represents one or more time-domain samples or transformed-domain samples of the uncorrelated signal, res represents one or more time-domain samples or transform-domain samples of the residual signal, u dmx,1 represents the downmix signal upmix parameter for the first output audio signal, and u dmx,2 represents the downmix signal upmix parameter for the second output audio signal, and u dec,1 represents the uncorrelated signal upmix parameter for the first output audio signal, and u dec,2represents the uncorrelated signal upmix parameter for the second output audio signal, max represents the maximum operator, and r represents a coefficient that describes the weighting of the uncorrelated signal according to the residual signal.
[0138] Upmix coefficient U dmx,1 ,U dmx,2 ,U dec,1 ,U dec,2 This is calculated with respect to the MPS2-1-2 parametric mode. For details, refer to the MPEG surround concept standard mentioned above.
[0139] In summary, embodiments of the present invention construct a concept that provides an output channel signal based on a downmix signal, a residual signal, and spatial data, with flexible adjustment of the weighting of the uncorrelated signals without any significant signaling overhead.
[0140] 7.5 Modified Embodiments
[0141] While several embodiments have been described in the context of apparatus, it is clear that these embodiments also represent descriptions of corresponding methods, where blocks or devices correspond to method steps or features of method steps. Similarly, embodiments described in the context of method steps also represent descriptions of corresponding blocks, items, or features of the corresponding apparatus. Some or all method steps can be performed by (or using) hardware devices, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such devices.
[0142] The encoded audio signal of the present invention can be stored in a digital storage medium or transmitted over a transmission medium (e.g., a wireless transmission medium or a wired transmission medium (e.g., the Internet)).
[0143] Depending on the specific implementation requirements, embodiments of the present invention may be implemented in hardware or in software. Implementation may be carried out using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, which has electronically readable control signals stored thereon and cooperates (or can cooperate) with a computer system programmable to perform each method. Therefore, the digital storage medium may be computer-readable.
[0144] Some embodiments of the present invention include a data carrier having electronically readable control signals and capable of cooperating with a programmable computer system on which one of the methods described in the present specification is performed.
[0145] Generally, embodiments of the present invention can be implemented as a computer program product having program code that is operable to perform one of the methods when the computer program product runs on a computer. The program code can be stored, for example, on a machine-readable carrier.
[0146] Other embodiments include a computer program stored in a machine-readable carrier for performing one of the methods described in this specification.
[0147] In other words, an embodiment of the present invention is a computer program having program code for executing one of the methods described in this specification when the computer program is running on a computer.
[0148] A further embodiment of the present invention is a data carrier (or digital storage medium or computer-readable medium) having a computer program recorded thereon that performs one of the methods described in this specification. The data carrier, digital storage medium or recording medium is generally tangible and / or non-transient.
[0149] A further embodiment of the present invention is a data stream or sequence of signals representing a computer program that performs one of the methods described in this specification. The data stream or sequence of signals can be configured to be transmitted, for example, over a data communication connection, such as the Internet.
[0150] Further embodiments include processing means, such as a computer or a programmable logic device, configured or adapted to perform one of the methods described in this specification.
[0151] Further embodiments include a computer on which a computer program is installed that performs one of the methods described in this specification.
[0152] Further embodiments of the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program that performs one of the methods described in the present specification to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server that transfers the computer program to the receiver.
[0153] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field-programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0154] The embodiments described above are merely illustrative of the principles of the present invention. Modifications and changes to the configurations and details described in this specification will be understood to be obvious to those skilled in the art. The present invention is therefore intended to be limited only by the immediate claims and not by the specific details provided by the descriptions and explanations of the embodiments in this specification.
[0155] 7.6 Further Embodiments
[0156] In the following, other embodiments of the present invention will be described with reference to Figure 8, which shows a schematic block diagram of a so-called hybrid residual decoder.
[0157] The hybrid residual decoder 800 shown in Figure 8 is very similar to the decoder 700 shown in Figure 7, and the above explanation is referred to. However, in the hybrid residual decoder 800, additional weighting (in addition to the application of upmix parameters) is applied only to the upmixed uncorrelated signal (which corresponds to signals 756 and 758 in decoder 700), and not to the upmixed residual signal (which corresponds to signals 760 and 762 in decoder 700). Thus, the weighting of the hybrid residual decoder 800 is somewhat simpler than the weighting in decoder 700, but it agrees well with the weighting by equation (14), for example.
[0158] The combined parametric and residual decoding (hybrid residual coding) shown in Figure 8 will be explained in some more detail below.
[0159] However, an overview is provided first.
[0160] In addition to using either a decorrelator-based mono-to-stereo upmix or residual coding as described in ISO / IEC 23003-3 (Section 7.11.1), hybrid residual coding allows for the dependent coupling of signals in both modes. As illustrated in Figure 8, the residual signal and the decorrelator output are mixed using time- and frequency-dependent weighting elements depending on the signal energy and spatial parameters.
[0161] The decryption process is described below.
[0162] TIFF0007862341000010.tif64169
[0163] The upmixing process is divided into the downmix, the uncorrelated output, and the residual. dmx It is calculated using the following formula. TIFF0007862341000011.tif17169
[0164] Upmixed uncorrelated output u dec It is calculated using the following formula. TIFF0007862341000012.tif16123
[0165] Upmixed residual signal u res It is calculated using the following formula. TIFF0007862341000013.tif16161
[0166] Energy E of the upmixed residual signal resThe energy Edec of the upmixed uncorrelated output is calculated as a sum over both the output channel ch and all time slots ts of one frame for each hybrid band, as follows: TIFF0007862341000014.tif22161
[0167] The upmixed uncorrelated output is obtained by calculating the following weighting element r for each hybrid band per frame: dec It is weighted using [this method]. TIFF0007862341000015.tif46164 Here, ε is a small number to prevent division by zero (e.g., ε = 1e-9 or 0 < ε <= 1e-5). However, in some embodiments, ε is set to zero ("E res <ε> to "E res It can be replaced with "=0".
[0168] All three upmix signals are added together to form the decoded output signal.
[0169] 8. Conclusion
[0170] In conclusion, the embodiments of the present invention are , remainder difference coding and parametric Combined coding To construct.
[0171] The present invention relates to a joint stereo based on the USAC integrated stereo tool. For coding parametric coding and residuals coding signal dependent Develop a coupling method. Instead of using a fixed residual bandwidth... ,workman ncoda The amount of residuals transmitted by Time and frequency variable I believe issue Depends on This is determined on the decoder side. teeth , Residue By mixing the difference signal and the uncorrelated output The required amount of decorrelation between output channels It is generated. In this way , corresponding audio encoding / decoding system by , to the encoded signal Dependent , at runtime, complete parametric coding and waveform preservation residual coding between blend can be done.
[0172] Embodiments according to the present invention are superior to conventional solutions. For example, in USAC, the MPEG surround 2-1-2 system , Department partial waveform preservation for transmits a band-limited or full-bandwidth residual signal Used for parametric stereo coding, or for integrated stereo. . When a band-limited residual is transmitted teeth , decorrelator of use did A parametric upmix is applied over the residual bandwidth. The drawback of this method is that the residual bandwidth is set to a fixed value at the encoder initialization.
[0173] In contrast, embodiments according to the present invention in are adaptable or parametric dependent to the signal of the residual bandwidth coding to Switching possible Nina be. Furthermore, when the downmix process in the parametric coding mode causes signal cancellation for a poor phase relationship, embodiments according to the present invention 、( for example, by providing an appropriate residual signal) Restoring the lost signal portion make it possible. The simplified downmix method Conventional parametric For coding does not cause signal cancellation more than MPS downmix but occur It is rare not It has been pointed out that . However, conventional simplified downmix cannot be used for partial waveform preservation where the residual signal is not defined in USAC in department , but embodiments according to the present invention There are can, inWaveform reconstruction (for example, partial waveform reconstruction is important) is important. It seems that... For the signal portion of (Selective partial waveform reconstruction) but Possible Nina ru.
[0174] In conclusion, embodiments of the present invention construct an apparatus, method, or computer program for audio encoding or decoding as described in the present specification.
Claims
1. A multichannel audio decoder (200;300;700;800) for providing at least two output audio signals (212,214;312,314;712,714) based on an encoded representation (210;310;710), The multi-channel audio decoder is configured to acquire one of the output audio signals based on an encoded representation of the downmix signal (222; 722), a plurality of encoded spatial parameters (726), and an encoded representation of the residual signal (226; 724). The multi-channel audio decoder is configured to fade between parametric decoding and residual decoding, depending on the residual signal and using weight adjustments that describe the contribution of the uncorrelated signal. Multi-channel audio decoder.
2. A multichannel audio encoder (100) for providing an encoded representation (112) of a multichannel audio signal (110), The multi-channel audio encoder is configured to acquire a downmix signal (122) based on the multi-channel audio signal, provide parameters (124) describing the inter-channel dependencies of the multi-channel audio signal, and provide a residual signal (126). The multi-channel audio encoder is configured to change the amount of residual signal included in the encoded representation depending on the multi-channel audio signal, The multi-channel audio encoder is configured to select the frequency bands included in the encoded representation of the residual signal, depending on the multi-channel audio signal. The multi-channel audio encoder is configured to selectively include the residual signal in the encoded representation for frequency bands in which the multi-channel audio signal is tonal. Multi-channel audio encoder.
3. The multichannel audio encoder according to claim 2, wherein the multichannel audio encoder is configured to change the bandwidth of the residual signal depending on the multichannel audio signal.
4. The multichannel audio encoder according to claim 2 or 3, wherein the multichannel audio encoder is configured to selectively include the residual signal in the encoded representation for time portions and / or frequency bands in which the signal components of the multichannel audio signal are canceled as a result of forming the downmix signal.
5. The multichannel audio encoder according to claim 4, wherein the multichannel audio encoder is configured to detect the cancellation of signal components of the multichannel audio signal in the downmix signal, and the multichannel audio encoder is configured to activate the provision of the residual signal in response to the result of the detection.
6. The multichannel audio encoder according to any one of claims 2 to 5, wherein the multichannel audio encoder is configured to calculate the residual signal using a linear combination of at least two channel signals of the multichannel audio signal and depending on the upmix coefficient used on the multichannel decoder side.
7. The multi-channel audio encoder is configured to determine and encode the upmix coefficient. Alternatively, the system is configured to derive the upmix coefficient from parameters describing the inter-channel dependency of the multi-channel audio signal. The multi-channel audio encoder according to claim 6.
8. The multichannel audio encoder according to any one of claims 2 to 7, wherein the multichannel audio encoder is configured to determine the amount of the residual signal included in the encoded representation as a time variable using a psychoacoustic model.
9. The multichannel audio encoder according to any one of claims 2 to 8, wherein the multichannel audio encoder is configured to determine the amount of the residual signal included in the encoded representation as a time variable, depending on the currently available bitrate.
10. A method (600) for providing at least two output audio signals based on an encoded representation, The step (610) includes obtaining one of the output audio signals based on an encoded representation of the downmix signal and a plurality of encoded spatial parameters and an encoded representation of the residual signal, A fade is performed between parametric decoding and residual decoding according to the residual signal and using weight adjustments that describe the contribution of the uncorrelated signal (620). method.
11. A method (400) for providing an encoded representation of a multichannel audio signal, The steps include: (410) acquiring a downmix signal based on the multi-channel audio signal; The steps include (420) providing parameters that describe the inter-channel dependencies of the multi-channel audio signal, Step (430) of providing a residual signal, Includes, The amount of residual signal included in the encoded representation changes depending on the multi-channel audio signal (440), The method includes the step of selecting frequency bands in which the residual signal is included in the encoded representation, depending on the multi-channel audio signal. The method includes the step of selectively including the residual signal in the encoded representation for frequency bands in which the multichannel audio signal is tonal. method.
12. A computer program for performing the method according to claim 10 or claim 11 when the computer program is running on a computer.
13. A multichannel audio decoder for providing at least two output audio signals based on an encoded representation, The multi-channel audio decoder is configured to acquire one of the at least two output audio signals based on an encoded representation of the downmix signal and a plurality of encoded spatial parameters and an encoded representation of the residual signal. The multichannel audio decoder is configured to fade between decoding using a scaled signal based on the downmix signal and residual-based decoding, depending on the residual signal and the uncorrelated signal acquired by the multichannel audio decoder, and using weight adjustments that describe the contribution of the uncorrelated signal. Multi-channel audio decoder.
14. A multichannel audio encoder (100) for providing an encoded representation (112) of a multichannel audio signal (110), The multi-channel audio encoder is configured to acquire a downmix signal (122) based on the multi-channel audio signal, provide parameters (124) describing the inter-channel dependencies of the multi-channel audio signal, and provide a residual signal (126). The multi-channel audio encoder is configured to change the amount of residual signal included in the encoded representation depending on the multi-channel audio signal, The multi-channel audio encoder is configured to selectively include the residual signal in the encoded representation for time portions and / or frequency bands in which the signal components of the multi-channel audio signal are canceled as a result of forming the downmix signal. The multi-channel audio encoder is configured to detect the cancellation of signal components of the multi-channel audio signal in the downmix signal, and the multi-channel audio encoder is configured to activate the provision of the residual signal in response to the result of the detection. Multi-channel audio encoder.
15. A multichannel audio decoder (200;300;700;800) for providing at least two output audio signals (212,214;312,314;712,714) based on an encoded representation (210;310;710), The multi-channel audio decoder is configured to acquire one of the output audio signals based on an encoded representation of the downmix signal (222; 722), a plurality of encoded spatial parameters (726), and an encoded representation of the residual signal (226; 724). The multi-channel audio decoder is configured to fade between parametric decoding and residual decoding, depending on the residual signal and using weight adjustments that describe the contribution of the uncorrelated signal. The multi-channel audio decoder is configured to perform a weighted combination (220; 780, 790, 792) of the downmix signal (222; 752, 754), the uncorrelated signal (224; 756, 758), and the residual signal (226; 760, 762; res) to obtain one of the output audio signals (212, 214; 712, 714). The multi-channel audio decoder, depending on the residual signal, uses weights (232; r; r) to describe the contribution of the uncorrelated signal in the weighted coupling. dec ) is configured to determine, The multi-channel audio decoder is configured to determine the weights that describe the contribution of the uncorrelated signal to the weighted coupling, depending on the uncorrelated signal. Multi-channel audio decoder.
16. A multichannel audio decoder for providing at least two output audio signals based on an encoded representation, The multi-channel audio decoder is configured to acquire one of the at least two output audio signals based on an encoded representation of the downmix signal and a plurality of encoded spatial parameters and an encoded representation of the residual signal. The multichannel audio decoder is configured to fade between decoding using a scaled signal based on the downmix signal and residual-based decoding, depending on the residual signal and the uncorrelated signal acquired by the multichannel audio decoder, and using weight adjustments that describe the contribution of the uncorrelated signal. The multi-channel audio decoder is configured to perform a weighted combination (220; 780, 790, 792) of the downmix signal (222; 752, 754), the uncorrelated signal (224; 756, 758), and the residual signal (226; 760, 762; res) to obtain one of the output audio signals (212, 214; 712, 714). The multi-channel audio decoder, depending on the residual signal, uses weights (232; r; r) to describe the contribution of the uncorrelated signal in the weighted coupling. dec ) is configured to determine, The multi-channel audio decoder is configured to determine the weights that describe the contribution of the uncorrelated signal to the weighted coupling, depending on the uncorrelated signal. Multi-channel audio decoder.
17. A method (600) for providing at least two output audio signals based on an encoded representation, The step (610) includes obtaining one of the output audio signals based on an encoded representation of the downmix signal and a plurality of encoded spatial parameters and an encoded representation of the residual signal, Depending on the residual signal, a fade is performed between parametric decoding and residual decoding (620). The method includes the step of performing a weighted combination (220; 780, 790, 792) of the downmix signal (222; 752, 754), the uncorrelated signal (224; 756, 758), and the residual signal (226; 760, 762; res) to obtain one of the output audio signals (212, 214; 712, 714), The method, depending on the residual signal, describes the contribution of the uncorrelated signal in the weighted combination using weights (232; r; r). dec This includes the step of determining, The method includes the step of determining the weights that describe the contribution of the uncorrelated signal to the weighted combination, depending on the uncorrelated signal. method.
18. A method for providing at least two output audio signals based on an encoded representation, The method includes the step of obtaining one of the at least two output audio signals based on an encoded representation of the downmix signal and a plurality of encoded spatial parameters and an encoded representation of the residual signal, The method includes a step of fading between decoding using a scaled signal based on the downmix signal and residual-based decoding, depending on the residual signal and the uncorrelated signal acquired by the multi-channel audio decoder, and using weight adjustments that describe the contribution of the uncorrelated signal. The method includes the step of performing a weighted combination (220; 780, 790, 792) of the downmix signal (222; 752, 754), the uncorrelated signal (224; 756, 758), and the residual signal (226; 760, 762; res) to obtain one of the output audio signals (212, 214; 712, 714), The method, depending on the residual signal, describes the contribution of the uncorrelated signal in the weighted combination using weights (232; r; r). dec This includes the step of determining, The method includes the step of determining the weights that describe the contribution of the uncorrelated signal to the weighted combination, depending on the uncorrelated signal. method.
19. A multichannel audio decoder (200;300;700;800) for providing at least two output audio signals (212,214;312,314;712,714) based on an encoded representation (210;310;710), The multi-channel audio decoder is configured to acquire one of the output audio signals based on an encoded representation of the downmix signal (222; 722), a plurality of encoded spatial parameters (726), and an encoded representation of the residual signal (226; 724). The multi-channel audio decoder is configured to fade between parametric decoding and residual decoding depending on the residual signal. The multi-channel audio decoder is configured to perform a weighted combination (220; 780, 790, 792) of the downmix signal (222; 752, 754), the uncorrelated signal (224; 756, 758), and the residual signal (226; 760, 762; res) to obtain one of the output audio signals (212, 214; 712, 714). The multi-channel audio decoder, depending on the residual signal, uses weights (232; r; r) to describe the contribution of the uncorrelated signal in the weighted coupling. dec ) is configured to determine, The multi-channel audio decoder is configured to determine the weights that describe the contribution of the uncorrelated signal to the weighted coupling, depending on the uncorrelated signal. The multi-channel audio decoder is configured to variably adjust the weights describing the contribution of the residual signal in the weighted coupling, depending on the uncorrelated signal. Multi-channel audio decoder.