Generation of a multi-channel audio signal

By combining a receiver, decorrelation unit, and envelope compensator with a trained artificial neural network, the distortion and complexity issues in generating stereo signals from mono submixed signals are solved, achieving more efficient multi-channel audio signal generation and improved audio quality.

CN122397079APending Publication Date: 2026-07-14KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480079814.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-12-11
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing technologies for generating stereo signals from mono downmixed signals suffer from distortion, artifacts, and audio quality degradation, and also have high data rates and processing complexity.

Method used

A receiver, decorrelator, envelope compensator, and output generator are used to generate multi-channel audio signals by decorrelating the frequency domain audio signal and compensating the time domain envelope using a trained artificial neural network.

Benefits of technology

It improves audio quality, reduces data rate and processing complexity, and provides more efficient audio signal reconstruction and an improved perceptual audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122397079A_ABST
    Figure CN122397079A_ABST
Patent Text Reader

Abstract

An apparatus comprising a receiver (101) arranged to receive a frequency domain audio signal and a decorrelator (107) generating a frequency domain decorrelated signal from the frequency domain audio signal. An envelope compensator (109) generates a compensated frequency domain decorrelated signal by applying a time domain envelope compensation to the decorrelated signal to reduce a difference between a time domain envelope of the frequency domain audio signal and a time domain envelope of the frequency domain decorrelated signal. The envelope compensator (109) comprises at least one trained artificial neural network having input values determined from at least one of the frequency domain audio signal and the frequency domain audio signal and generating time domain envelope compensation data for the time domain envelope compensation. An output generator (105) generates a multi-channel audio signal by upmixing from the frequency domain audio signal and the frequency domain decorrelated signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the generation of multichannel audio signals, and more particularly (but not exclusively) to the generation of stereo signals from the upmixing of a mono downmix signal. Background Technology

[0002] Spatial audio applications have become numerous and widespread, increasingly forming at least part of many audiovisual experiences. In fact, the continuous development of new and improved spatial experiences and applications has led to an increased demand for audio processing and rendering.

[0003] For example, virtual reality (VR) and augmented reality (AR) have received increasing attention in recent years, with many implementations and applications entering the consumer market. In fact, devices are being developed for both rendering experiences and capturing or recording suitable data for such applications. For instance, relatively low-cost devices are being developed to allow game consoles to deliver a full VR experience. This trend is expected to continue, and its pace will indeed accelerate as the VR and AR markets reach a considerable size in the near future. In the audio field, a prominent area of ​​exploration is the reproduction and synthesis of real-world and natural spatial audio. The ideal goal is to produce natural audio sources so that users cannot distinguish between synthesized and original audio sources.

[0004] Much research and development has focused on providing efficient and high-quality audio coding and decoding for spatial audio. A common spatial audio representation is multichannel audio representation, including stereo representation, and efficient coding methods have been developed based on downmixing the multichannel audio signal to a sub-mixer channel with fewer channels. One of the major advances in low bit-rate audio coding is the use of parametric multichannel decoding, where the sub-mixer signal is generated along with parametric data, which can be used to upmix the sub-mixer signal to reconstruct the multichannel audio signal.

[0005] Specifically, instead of traditional center-side or intensity decoding, parametric multichannel audio decoding uses a multichannel input signal to downmix a smaller number of channels (e.g., two to one) and extracts multichannel picture (stereo) parameters. The downmixed signal is then encoded using a more conventional audio decoder (e.g., a mono audio encoder). The downmixed data is combined with the encoded multichannel parameter data to generate a suitable audio bitstream. This bitstream is then sent to a decoder, where the process is reversed. The downmixed audio signal is first decoded, and then the reconstructed multichannel audio signal is guided by the encoded multichannel picture upmix parameters.

[0006] An example of stereo decoding is described in the following document: E. Schuijers, W. Oomen, B. den Brinker, J. Breebaart, “Advances in High-Quality Audio Parametric Decoding,” 114th AES Conference, Amsterdam, Netherlands, 2003, Preprint 5852. In the described method, the downmixed mono signal is parameterized by utilizing the natural separation of the signal into three components (objects): transient, sine wave, and noise. Further details are provided in the following document describing how parametric stereo can be implemented with low (decoder) complexity when combined with spectral band replication (SBR): E. Schuijers, J. Breebaart, H. Purnhagen, J. Engdegård, “Low-Complexity Parametric Stereo Decoding,” 116th AES Conference, Berlin, Germany, 2004, Preprint 6073.

[0007] In the described method, the decoding-side multichannel reproduction is based on the use of a so-called decorrelation process. The decorrelation process generates a decorrelation auxiliary signal from the downmix signal (i.e., from the parametric stereo mono signal). Specifically, the decorrelation signal is used during the upmix process to control and reconstruct the coherence (stereo width in the case of stereo) between / within the different channels. During stereo reconstruction, both the mono signal and the decorrelation auxiliary signal are used to generate an upmixed stereo signal based on the upmix parameters. Specifically, the two signals can be multiplied by a 2x2 matrix with time- and frequency-dependent coefficients determined according to the upmix parameters to provide the output stereo signal.

[0008] However, while parametric stereo (PS) and similar downmixing encoding / decoding methods represent a leap forward from traditional stereo and multichannel decoding, they are not optimal in all cases. In particular, known methods tend to introduce distortions, variations, artifacts, etc., which can create discrepancies between the (raw) multichannel audio signal input to the encoder and the multichannel audio information reproduced at the decoder. Typically, audio quality may degrade, and multichannel reproduction may be imperfect. Furthermore, the data rate may still be higher than desired, and / or the complexity / resource usage of the processing involved may exceed preferred values. In particular, generating a decorrelation auxiliary signal to obtain optimal audio quality for the upmixed multichannel signal is challenging. For example, decorrelation processes are known to gradually eliminate transients over time and may introduce additional artifacts, such as back echoes due to the filter characteristics of the decorrelator.

[0009] Therefore, improved approaches will be advantageous. In particular, approaches that allow for increased flexibility, improved adaptability, improved performance, increased audio quality, improved audio quality with data rate trade-offs, reduced complexity and / or resource usage, reduced computational load, facilitated implementation, and / or improved spatial audio experience will be advantageous. Summary of the Invention

[0010] Therefore, the present invention attempts to alleviate, reduce or eliminate one or more of the above-mentioned disadvantages, either alone or in any combination.

[0011] According to one aspect of the present invention, an apparatus for generating a multi-channel audio signal is provided, the apparatus comprising: a receiver arranged to receive a frequency-domain audio signal; a decorrelation unit arranged to generate a frequency-domain decorrelation signal by applying decorrelation to the frequency-domain audio signal; an envelope compensator arranged to generate an envelope-compensated frequency-domain decorrelation signal by applying time-domain envelope compensation to the frequency-domain decorrelation signal, the time-domain envelope compensation reducing the difference between the time-domain envelope of the frequency-domain audio signal and the time-domain envelope of the frequency-domain decorrelation signal, the envelope compensator including at least one trained artificial neural network having input values ​​determined based on at least one of the frequency-domain audio signal and the frequency-domain decorrelation signal and generating time-domain envelope compensation data for the time-domain envelope compensation; and an output generator arranged to generate the multi-channel audio signal by upmixing from the frequency-domain audio signal and the envelope-compensated frequency-domain decorrelation signal.

[0012] In many embodiments, this method can provide an improved audio experience. For many signals and scenarios, this method can provide improved generation / reconstruction of multichannel audio signals with improved perceived audio quality. This method can provide a particularly advantageous arrangement that can facilitate and / or improve the possibility of utilizing artificial neural networks in audio processing, including conventional audio encoding and / or decoding, in many embodiments and scenarios. This method can allow for the advantageous use of artificial neural networks in generating multichannel audio signals from downmixed audio signals.

[0013] This method can provide an efficient implementation and, in many embodiments, can allow for reduced complexity and / or resource usage.

[0014] This method provides a particularly efficient and high-performance approach for improving and increasing the correspondence of properties between the received audio signal and the auxiliary decorrelation signal used to generate the multichannel audio signal. Specifically, this method can effectively mitigate and reduce the effects of the generated decorrelation signal, which does not possess ideal temporal properties, particularly errors or deviations in the signal level of the decorrelation signal.

[0015] In many embodiments, the frequency domain audio signal can be a mono audio signal and / or the multi-channel audio signal can be a stereo audio signal. In many embodiments, the receiver can also receive upmixing parameters, and the upmixing can depend on the upmixing parameters.

[0016] Upmixing parameter data may include parameters (values) that correlate the properties of the downmixed signal with the properties of the multichannel audio signal. Upmixing parameter data may include data indicating the relative properties between the channels of the multichannel audio signal. Upmixing parameter data may include data indicating the differences in properties between the channels of the multichannel audio signal. Upmixing parameter data may include data perceptually related to the synthesis of the multichannel audio signal. Properties may be, for example, differences in phase and / or intensity and / or timing and / or correlation. In some embodiments and scenarios, upmixing parameter data may represent abstract properties that are not directly understood by humans / experts (but often facilitate better reconstruction / lower data rates, etc.). Upmixing parameter data may include data on at least one of the following: inter-channel intensity difference, inter-channel timing difference, inter-channel correlation, and / or inter-channel phase difference of the channels comprising the multichannel audio signal.

[0017] The artificial neural network can be a trained artificial neural network trained on training data, which includes a training downmixed audio signal generated from a training multichannel audio signal and training upmixing parameter data. Training employs a cost function that compares the training multichannel audio signal with an upmixed multichannel signal generated by an audio device from the training downmixed signal and from an envelope-compensated frequency domain decorrelation signal. The artificial neural network can also be a trained artificial neural network trained on training data, which includes training data representing a range of relevant audio sources, including recordings of video, film, telecommunications, etc.

[0018] An artificial neural network can be a trained artificial neural network trained with training data and using a cost function, the training data having training input data including training audio signals, the cost function including a contribution indicating the difference between the generated multichannel audio signal and the training input multichannel audio signal, and / or the cost function indicating the difference between the time-domain envelope properties of the frequency-domain audio signal and the compensated frequency-domain decorrelation signal.

[0019] The generator can be configured to generate a multichannel audio signal by applying matrix multiplication to a compensated frequency-domain decorrelation signal and a frequency-domain audio signal, where the coefficients of the matrix are determined as a function of the parameters of the upmixing parameter data. The matrix can be time- and frequency-dependent.

[0020] Specifically, the audio device can be an audio decoder device.

[0021] The receiver can receive frequency-domain audio signals from any suitable internal or external source. For example, an audio device may include a bitstream receiver for receiving a bitstream comprising a time-domain audio signal (particularly a mono signal). The bitstream receiver can extract the time-domain audio signal and transform it to the frequency domain to generate a frequency-domain audio signal, which can then be forwarded to a receiver configured to receive the frequency-domain audio signal. The bitstream receiver can also receive additional data in the bitstream, such as upmixed data, location information, acoustic environmental data, etc.

[0022] A trained artificial neural network may include input nodes that receive samples of frequency-domain audio signals and / or data values / features derived therefrom. A trained artificial neural network may include input nodes that receive samples of frequency-domain decorrelation signals and / or data values / features derived therefrom. A trained artificial neural network may include output nodes that generate samples of compensated frequency-domain decorrelation signals and / or can generate data from which compensated frequency-domain decorrelation signals can be generated (e.g., using analysis / predetermined functions and / or without employing an artificial neural network).

[0023] According to an optional feature of the invention, at least one trained artificial neural network includes a trained artificial neural network arranged to receive an input value determined based on a frequency domain audio signal and to generate time-domain envelope compensation data indicating / representing / as a time-domain envelope estimate ( / signal) of the frequency domain audio signal, wherein an envelope compensator is arranged to generate an envelope-compensated frequency-domain decorrelation signal based on the time-domain envelope estimate ( / signal) of the frequency domain audio signal.

[0024] This can provide particularly efficient and high-performance operation in many scenarios, and can often lead to a significant improvement in the audio quality of the generated multichannel audio signal.

[0025] In some embodiments, the time-domain envelope estimate can be determined as a signal, which may be, for example, a time-domain signal representing a time-domain sample of the envelope. In other embodiments, the time-domain envelope estimate signal may be determined and represented by frequency-domain samples of the signal. In some embodiments, the time-domain envelope estimate may alternatively or additionally be represented by one or more parameters / characteristics / attributes of the time-domain envelope. For example, the time-domain envelope may be represented by features or parameters, including, for example, average level, intermediate level, maximum level or minimum level, a measure of change such as median or average frequency, etc.

[0026] An envelope compensator can be configured to use an artificial neural network to determine a time-domain envelope estimate (signal) of a frequency-domain audio signal; compare a time-domain envelope estimate (signal) of a frequency-domain decorrelated signal with a time-domain envelope estimate (signal) of the frequency-domain audio signal; and compensate the frequency-domain decorrelated signal based on the comparison / difference between these time-domain envelope estimates (signals).

[0027] According to an optional feature of the invention, at least one trained artificial neural network includes a trained artificial neural network arranged to receive an input value determined based on a frequency-domain decorrelation signal, and to generate time-domain envelope compensation data indicating / representing / as a time-domain envelope estimate ( / signal) of the frequency-domain decorrelation signal, wherein an envelope compensator is arranged to generate an envelope-compensated frequency-domain decorrelation signal based on the time-domain envelope estimate ( / signal) of the frequency-domain decorrelation signal.

[0028] This can provide particularly efficient implementation and / or improved performance. In many scenarios, improved audio quality of multi-channel output audio signals can be achieved.

[0029] An envelope compensator can be configured to use an artificial neural network to determine a time-domain envelope estimate (signal) of a frequency-domain decorrelated signal; compare the time-domain envelope estimate (signal) of the frequency-domain decorrelated signal with a time-domain envelope estimate (signal) of a frequency-domain audio signal; and compensate the frequency-domain decorrelated signal based on the comparison / difference between these time-domain envelope estimates (signals).

[0030] In some embodiments, the time-domain envelope estimate can be determined as a signal, which may be, for example, a time-domain signal representing a time-domain sample of the envelope. In other embodiments, the time-domain envelope estimate signal may be determined and represented by frequency-domain samples of the signal. In some embodiments, the time-domain envelope estimate may alternatively or additionally be represented by one or more parameters / characteristics / attributes of the time-domain envelope. For example, the time-domain envelope may be represented by features or parameters, including, for example, average level, intermediate level, maximum level or minimum level, a measure of change such as median or average frequency, etc.

[0031] According to an optional feature of the invention, the envelope compensator includes: a first envelope estimator arranged to determine a time-domain envelope estimate ( / signal) of a frequency-domain audio signal; a second envelope estimator arranged to determine a time-domain envelope estimate ( / signal) of a frequency-domain decorrelated signal; and wherein at least one trained artificial neural network includes a trained artificial neural network arranged to receive an input value determined based on the time-domain envelope estimate ( / signal) of the frequency-domain audio signal and an input value determined based on the time-domain envelope estimate (signal) of the frequency-domain decorrelated signal, and to generate time-domain envelope-compensated data including samples of the envelope-compensated frequency-domain decorrelated signal (frequency-domain or time-domain).

[0032] This can provide particularly efficient implementation and / or improved performance. In many scenarios, improved audio quality of multi-channel output audio signals can be achieved.

[0033] In many scenarios, particularly advantageous operation and performance are achieved by using a first envelope estimator and / or a second envelope estimator implemented using a trained artificial neural network as described above.

[0034] In some embodiments, the time-domain envelope estimate can be determined as a signal, which may be, for example, a time-domain signal representing a time-domain sample of the envelope. In other embodiments, the time-domain envelope estimate signal can be determined and represented by frequency-domain samples of the signal. In some embodiments, the time-domain envelope estimate signal may alternatively or additionally be represented by one or more parameters / characteristics / attributes of the time-domain envelope. For example, the time-domain envelope may be represented by features or parameters, including, for example, average level, intermediate level, maximum level or minimum level, a measure of change such as median or average frequency, etc.

[0035] According to an optional feature of the invention, at least one of the trained artificial neural networks is arranged to receive an input value determined according to a frequency domain audio signal and an input value determined according to a frequency domain decorrelation signal, and to generate time-domain envelope-compensated data to include (frequency domain or time domain) samples of the envelope-compensated frequency domain decorrelation signal.

[0036] This can provide particularly efficient implementation and / or improved performance. In many scenarios, it can achieve improved audio quality for multi-channel output audio signals. In many scenarios, it can provide a highly computationally efficient implementation that can provide highly accurate compensation for a wide range of audio signals.

[0037] A trained artificial neural network can specifically include input nodes for receiving samples of frequency domain audio signals and frequency domain decorrelation signals.

[0038] According to an optional feature of the invention, at least one trained artificial neural network is arranged to generate temporal envelope compensation data for a first time interval based on input values ​​determined according to samples of at least one of a frequency-domain audio signal and a frequency-domain decorrelation signal in a second time interval, the second time interval having a duration not less than twice that of the first time interval.

[0039] This can provide particularly efficient implementation and / or improved performance. In many scenarios, improved audio quality of multi-channel output audio signals can be achieved.

[0040] The trained artificial neural network may include input nodes for receiving samples of frequency-domain audio signals and / or frequency-domain decorrelation signals in a time interval, the duration of which exceeds the duration of the time interval in which the output nodes of the trained artificial neural network generate time-domain envelope compensation data.

[0041] The time-domain envelope compensation data can be samples of the frequency-domain decorrelation signal.

[0042] In some embodiments, the envelope compensator may be arranged to generate frequency samples of the envelope-compensated frequency-domain decorrelation signal for a first time interval based on the output of the trained artificial neural network, when the trained artificial neural network has input values ​​determined according to samples of at least one of the frequency-domain audio signals and the frequency-domain audio signal in a second time interval, the second time interval having a duration not less than twice that of the first time interval.

[0043] According to optional features of the invention, the envelope compensator includes: a first feature determiner arranged to determine a first feature set based on a frequency domain decorrelation signal; a second feature determiner arranged to determine a second feature set based on a frequency domain audio signal; wherein at least one trained artificial neural network includes a trained artificial neural network arranged to generate a feature scaling factor set, the artificial neural network having inputs including the first feature set and the second feature set; and the envelope compensator further includes: circuitry arranged to apply scaling factors to the first feature set to generate a first compensated feature set; and a feature decoder for generating an envelope-compensated frequency domain decorrelation signal from the first compensated feature set.

[0044] This can provide particularly efficient implementation and / or improved performance. Improved audio quality of multi-channel output audio signals can be achieved in many scenarios. This arrangement (especially scaling based on artificial neural networks and performed in the feature domain rather than the signal domain) can provide highly advantageous operation and performance in many embodiments and scenarios. It has been found that this is not only highly computationally efficient for many signals, but also results in a significant improvement in the audio quality of multi-channel audio signals.

[0045] The first and / or second feature sets can be learned by a corresponding artificial neural network for the frequency-domain audio signal and the frequency-domain decorrelation signal. The artificial neural network can learn time-domain envelope information, but not in the frequency domain. This means that features may not always be easily interpretable and may include features jointly learned across frequency bands and time (spectral time).

[0046] According to an optional feature of the invention, the feature decoder includes a trained artificial neural network having an input node that receives a first set of compensated features and an output node that provides samples of envelope-compensated frequency-domain decorrelation signals.

[0047] This can provide particularly efficient implementation and / or improved performance in many scenarios.

[0048] According to an optional feature of the invention, the first feature determiner includes an artificial neural network having an input node that receives samples of frequency-domain decorrelation signals and an output node that provides a first set of features.

[0049] This can provide particularly efficient implementation and / or improved performance in many scenarios.

[0050] According to an optional feature of the invention, the second feature determiner includes an artificial neural network having an input node that receives samples of frequency-domain audio signals and an output node that provides a second set of features.

[0051] This can provide particularly efficient implementation and / or improved performance in many scenarios.

[0052] According to an optional feature of the invention, at least one artificial neural network is trained using training data comprising a number of training frequency-domain audio signals and a cost function, the cost function depending on a measure of the difference between the time-domain envelope of the training frequency-domain audio signals and the time-domain envelope of a compensated frequency-domain decorrelation signal generated for the training frequency-domain audio signals.

[0053] This can provide particularly efficient implementation and / or improved performance. In many scenarios, improved audio quality of multi-channel output audio signals can be achieved. Envelope compensators can be efficiently trained based on such a training process to provide improved envelope compensation.

[0054] According to an optional feature of the invention, at least one trained artificial neural network includes an input convolutional layer that receives samples of frequency domain audio signals and frequency domain decorrelation signals, and at least one fully connected hidden layer.

[0055] This can provide particularly advantageous and effective envelope compensation.

[0056] According to an optional feature of the invention, the envelope compensator is arranged to generate an envelope-compensated frequency-domain decorrelation signal by applying time-domain envelope compensation only to a subset of the higher-frequency subbands of the frequency-domain decorrelation signal.

[0057] This can provide particularly efficient implementation and / or improved performance in many scenarios.

[0058] According to one aspect of the present invention, a method for generating a multi-channel audio signal is provided, the method comprising: receiving a frequency-domain audio signal; generating a frequency-domain decorrelation signal by applying decorrelation to the frequency-domain audio signal; generating an envelope-compensated frequency-domain decorrelation signal by applying time-domain envelope compensation to the frequency-domain decorrelation signal, the time-domain envelope compensation reducing the difference between the time-domain envelope of the frequency-domain audio signal and the time-domain envelope of the frequency-domain decorrelation signal, the time-domain envelope compensation including at least one trained artificial neural network having input values ​​determined based on at least one of the frequency-domain audio signal and the frequency-domain decorrelation signal and generating time-domain envelope compensation data for the time-domain envelope compensation; and generating the multi-channel audio signal by upmixing from the frequency-domain audio signal and the envelope-compensated frequency-domain decorrelation signal.

[0059] These and other aspects, features, and advantages of the invention will become apparent and will be clarified with reference to the embodiments described below. Attached Figure Description

[0060] Embodiments of the invention will be described by way of example only with reference to the accompanying drawings, wherein:

[0061] Figure 1 Some elements of an example audio device according to some embodiments of the present invention are shown;

[0062] Figure 2 Some elements of an example time-frequency converter for an audio device according to some embodiments of the present invention are shown;

[0063] Figure 3 An example of the structure of an artificial neural network is shown;

[0064] Figure 4 An example of a node in an artificial neural network is shown;

[0065] Figure 5 Some elements of an example of an envelope compensator for an audio device according to some embodiments of the present invention are shown;

[0066] Figure 6 Some elements of an example of an envelope compensator for an audio device according to some embodiments of the present invention are shown;

[0067] Figure 7 An example of an audio signal and a decorrelation signal generated from the audio signal is shown;

[0068] Figure 8 Some elements of an example of an envelope compensator for an audio device according to some embodiments of the present invention are shown;

[0069] Figure 9 Some elements of an example of an envelope compensator for an audio device according to some embodiments of the present invention are shown;

[0070] Figure 10 Some elements of an example audio device according to some embodiments of the present invention are shown;

[0071] Figure 11 An example of the arrangement of elements for training a neural network is shown; and

[0072] Figure 12 Some elements of a processor for implementing an audio device are shown in some embodiments of the present invention. Detailed Implementation

[0073] Figure 1 Some components of an audio device arranged to generate multi-channel audio signals are shown.

[0074] The audio device includes a receiver 101 arranged to receive a data signal / bitstream comprising an audio signal, specifically a downmixed audio signal, which is a downmix of a multichannel audio signal recreated / generated by the audio device. The following description will focus on the case where the multichannel audio signal is a stereo signal and the downmixed signal is a mono signal; however, it should be understood that the described methods and principles are equally applicable to multichannel audio signals having more than two channels and downmixed signals having more than one channel (although fewer channels than the multichannel audio signal).

[0075] The received audio signal is a frequency-domain audio signal, and in some embodiments, the frequency-domain audio signal is received directly from a remote source that can generate a bitstream including a frequency representation of the audio signal. In other embodiments, the remote source can generate a time-domain representation of the audio signal, and a local frequency converter can be arranged to generate a (time-)frequency representation of the time-domain representation. Therefore, in some embodiments, receiver 101 can receive a frequency-domain audio signal from an internal source such as an internal time-to-frequency converter.

[0076] In particular, Figure 1 In the example, the audio device includes a filter bank 103 arranged to generate a frequency sub-band representation of the received time-domain mixed audio signal. Typically, the audio device may include a filter bank 103 applied to the audio signal such that it is divided into frequency sub-bands.

[0077] The filter bank can be a group of quadrature mirror filters (QMFs), or it can be implemented, for example, by a fast Fourier transform (FFT). However, it should be understood that many other filter banks and methods for dividing an audio signal into multiple sub-band signals are known and can be used. Specifically, the filter bank can be a group of complex-valued pseudo-QMFs, thereby producing, for example, 32 or 64 complex-valued sub-band signals.

[0078] Furthermore, processing is typically performed in time periods or time slots. In most embodiments, the audio signal is divided into time intervals / segments and transformed to the frequency domain / subband domain by applying, for example, FFT or QMF filtering to samples of each signal. For example, each channel of the downmixed audio signal can be divided into time intervals of, for example, 2048, 1024, or 512 samples. These signals can then be processed to generate, for example, 64, 32, or 16 subband samples. Thus, a sample set can be determined for each subband of the downmixed audio signal.

[0079] It should be noted that the number of time-domain samples is not directly related to the number of subbands. Typically, for a so-called N-band critical sampling filter bank, every N input samples will result in N subband samples (one per subband). An oversampling filter bank will produce more output samples. For example, for every N input samples, it will generate k... N output samples, that is, k consecutive samples for each frequency band.

[0080] In some embodiments, subbands are generated to have the same bandwidth, but in other embodiments, subbands are generated to have different bandwidths, for example, reflecting the sensitivity of human hearing to different frequencies.

[0081] Figure 2 An example of a method for generating different bandwidths using a hybrid analysis / synthetic filter bank approach is shown.

[0082] In this example, this is achieved through a combination of a complex-valued pseudo-QMF group 201 for the lower frequency band and a small filter group 203 to achieve higher frequency resolution, as desired by binaural perception in the human auditory system. The result is a hybrid filter group with a logarithmic filter band center-frequency spacing that follows human perception similar to the equivalent rectangular bandwidth (ERB). To compensate for the filtering delay of the small filter group 203, a delay 205 is introduced for the higher frequency subbands.

[0083] In a specific example, the time-domain signal Through having The signal is fed by downsampled complex exponential modulation (QMF) groups in each frequency band. 64 time-domain samples. Each frame results in QMF samples In the time slot One time slot, in which The lower frequency bands are then further split and filtered using an additional complex modulation filter bank. Higher time slots are delayed to ensure that the filtered signal from the lower frequency bands is synchronized with the higher frequency bands, as filtering introduces the delay. This ultimately results in a structure where, for every 64 time-domain samples… In the time slot Mixed QMF samples are generated at the location A time slot m, where For example, the total number of mixed frequency bands .

[0084] The received data signal also includes upmixing parameter data for upmixing the downmixed audio signal. The upmixing parameter data can specifically be a set of parameters indicating the relationship between signals of different audio channels of a multichannel audio signal (specifically, a stereo signal) and / or the relationship between the downmixed signal and the audio channels of the multichannel audio signal. Typically, upmixing parameters can indicate time difference, phase difference, level / intensity difference, and / or similarity measures such as correlation. Typically, upmixing parameters are provided on a per-time and per-frequency basis (time-frequency slice). For example, new parameters can be provided periodically for a set of sub-bands. Parameters can specifically include inter-channel phase difference (IPD), overall phase difference (OPD), inter-channel correlation (ICC), and channel phase difference (CPD) parameters, as known from the parameter stereo encoding (and from the higher channel encoding).

[0085] Typically, the downmixed audio signal is encoded, and receiver 101 may include a decoder function that decodes the downmixed audio signal (i.e., the mono signal in a particular example).

[0086] Receiver 101 is coupled to generator 105, which generates a multi-channel audio signal from a frequency-domain audio signal. In this example, generator 105 is arranged to generate the multi-channel audio signal by upmixing (at least) the frequency-domain audio signal and an auxiliary audio signal generated from the frequency-domain audio signal. Upmixing is performed according to parameter upmixing data.

[0087] Specifically, for stereo, generator 105 can generate an output multichannel audio signal by applying 2x2 matrix multiplication to samples of the downmixed audio signal and the auxiliary audio signal. The coefficients of the 2x2 matrix are typically determined based on the upmixing parameters according to the upmixing parameter data, using both time and frequency band. For other upmixing operations, such as converting a mono or stereo downmixed signal to a five-channel multichannel audio signal, generator 105 can apply matrix multiplication with a matrix of appropriate dimensions.

[0088] It will be understood that many different methods for generating such multichannel audio signals from downmixed and auxiliary audio signals and for determining suitable matrix coefficients from upmixing parameter data will be known to those skilled in the art, and any suitable method may be used. Specifically, various methods for parametric stereo upmixing based on downmixed and auxiliary audio signals are well known to those skilled in the art.

[0089] An auxiliary audio signal is generated from a frequency-domain audio signal by applying a decorrelation operation / function to the frequency-domain audio signal, and thus the auxiliary audio signal is the decorrelated signal corresponding to the frequency-domain audio signal. It has been found that by generating the decorrelated signal and mixing it with the downmixed audio signal (specifically, a mono audio signal) for stereo upmixing, an improved quality of the upmixed signal is perceived, and many decoders have been developed to take advantage of this. The decorrelated signal can be generated, for example, by employing a decorrelator in the form of a full-phase filter applied to the downmixed audio signal.

[0090] therefore, Figure 1 The audio device includes a decorreductor 107, which is arranged to generate a frequency-domain decorreducted signal by applying a decorreductance function / process to the frequency-domain audio signal.

[0091] It should be understood that many different methods for generating frequency-domain decorrelation signals from frequency-domain audio signals are known, and any suitable method may be used without departing from the present invention.

[0092] For example, decorreductor 107 may include filtering the input signal with a decaying noise sequence, which is not much different from a reverberator. The decay characteristic may be frequency-dependent. As another example, decorreductor may include an all-pass filter bank that applies an appropriate frequency-dependent phase shift to the signal. Another time-domain approach involves dividing the signal into overlapping blocks and shuffling them before windowing and resynthesizing them. However, the latter tends to work well only for random signals such as fixed noise.

[0093] However, while it has indeed been found that using decorrelation signals in upmixing can provide improved audio quality, Figure 1 The audio device is arranged not to directly use the generated frequency domain decorrelation signal. Instead, decorrelation unit 107 is coupled to envelope compensator 109, which is arranged to perform time-domain envelope compensation of the frequency domain decorrelation signal to generate an envelope-compensated frequency domain decorrelation signal, hereinafter referred to as the compensated frequency domain decorrelation signal for simplicity. Envelope compensator 109 is coupled to generator 105, feeding the compensated frequency domain decorrelation signal to generator 105, which is then used as an auxiliary audio signal for upmixing. Thus, generator 105 generates a multi-channel audio signal by upmixing the frequency domain audio signal and the compensated frequency domain decorrelation signal.

[0094] Envelope compensator 109 is configured to apply time-domain envelope compensation to the frequency-domain decorrelated signal. Time-domain envelope compensation reduces the envelope difference between the time-domain audio signal and the time-domain decorrelated signal. Therefore, if all signals are converted to the time domain, the envelope of the compensated decorrelated signal will match the envelope of the downmixed audio signal more closely than the envelope of the uncompensated decorrelated signal. In other words, the difference between the time-domain envelope of the decorrelated signal and the time-domain envelope of the downmixed audio signal at the output of envelope compensator 109 is less than the difference between the time-domain envelopes of the time-domain decorrelated signal and the downmixed audio signal at the input of envelope compensator 109.

[0095] The time-domain envelope of a signal can specifically be a smoothed representation of the extrema of the underlying signal. It is the boundary into which the signal is contained. Known methods for deriving the time-domain envelope from a signal include wave rectifiers, which produce the absolute value of the signal, and then extracting a smoothed line through the successive peaks of the signal. In some cases, the envelope is allowed to shift rapidly upwards to properly follow transients, while it is configured to decay slowly.

[0096] The time-domain envelope of an audio signal can be a time-domain signal corresponding to / to be generated by a low-pass filter on the signal or a rectified version of the signal (i.e., the signal produced by rectifying the original signal). The low-pass filter can be an ideal low-pass filter that can have a cutoff frequency of, for example, 5 Hz, 10 Hz, 20 Hz, 50 Hz, or 100 Hz. The time-domain envelope can also be referred to as the time-domain amplitude or time-domain level.

[0097] The envelope of an audio signal can reflect slow changes in signal level, thus removing or filtering out rapid / high changes. The envelope can also reflect (smooth) signal power / energy levels.

[0098] The envelope of a signal can be considered as a positive (i.e., no zero-crossing) signal that follows the slowly varying amplitude dynamics of the signal. The envelope can be considered as being generated by (half-wave or full-wave) rectification followed by low-pass filtering (e.g., with a cutoff frequency as described above). The envelope can also be the amplitude of an analytical signal (with negative frequency terms removed from the true signal).

[0099] The inventors have recognized that a problem with many decorrelation processes is that the decorrelated signal will not ideally reflect the properties of the original downmixed audio signal that generated it, and in particular, it will not have the same envelope / signal level / amplitude / power characteristics. The inventors have also recognized that this effect can be mitigated and reduced by implementing specific applications of artificial neural networks.

[0100] exist Figure 1In the audio device, a two-stage process is employed, whereby first a decorrelated signal is generated from the audio signal, and then compensation is performed using a neural network method. Specifically, the generation of multi-channel audio signals is improved through a specific arrangement, wherein a trained artificial neural network is included in a specific functional block that performs temporal envelope compensation for the generated decorrelated signal used to upmix the multi-channel audio signal. The inventors have recognized that practical implementation and high performance can be achieved through a specific structure in which the trained artificial neural network is specifically trained as part of the temporal envelope compensation of the frequency domain decorrelated signal. In particular, compensation is performed on the frequency domain signal, but with temporal characteristics. The trained artificial neural network is particularly efficient in this case, and can be trained efficiently using relatively low-complexity training signals and processes.

[0101] The artificial neural network used in the described function can be a network of nodes arranged in layers, with each node retaining its own node value. Figure 3 An example of a portion of an artificial neural network is shown.

[0102] The node value of a given node can be calculated to include contributions from some or typically all nodes in previous layers of the artificial neural network. Specifically, the node value of a node can be calculated as a weighted sum of the node values ​​of the outputs of all nodes in the previous layers. Typically, biases can be added, and the result can be subjected to an activation function. Activation functions provide the necessary portion of each neuron by typically providing nonlinearity. This nonlinearity and the activation function have a significant impact on the learning and adaptation process of the neural network. Therefore, node values ​​are generated based on the node values ​​of previous layers.

[0103] The artificial neural network may specifically include an input layer 301, which includes multiple nodes that receive input data values ​​from the artificial neural network. Therefore, the node values ​​of the nodes in the input layer can typically be directly the input data values ​​of the artificial neural network, and thus can be derived without being calculated from other node values.

[0104] Artificial neural networks may also include zero, one or more hidden layers 303, 305, or processing layers. For each of these layers, node values ​​are typically generated based on the node values ​​of the nodes in the previous layers, and specifically, based on a weighted combination and added bias, followed by an activation function (e.g., a sigmoid, ReLU, or Tanh function may be applied).

[0105] Specifically, such as Figure 4 As shown, each node, also known as a neuron, can receive input values ​​(from nodes in previous layers) and thereby compute node values ​​based on these values. Typically, this involves first generating values ​​as linear combinations of input values, where each of these input values ​​is weighted by weights: Where w refers to the weight, x refers to the node in the previous layer, and n refers to the index of the different node in the previous layer.

[0106] The activation function can then be applied to the resulting combination. For example, the node value l can be determined as: This function can be, for example, a rectified linear unit function (as described in the following document: Xavier Glorot, Antoine Bordes, Yoshua, Proceedings of the 14th International Conference on Artificial Intelligence and Statistics, PMLR 15:315-323, 2011):

[0107] Other commonly used functions include the sigmoid function or the tanh function. In many implementations, multiple functions can be used to compute the node output or value. For example, the ReLU and sigmoid functions can be combined using activation functions, such as:

[0108] Such operations can be performed by each node of an artificial neural network (typically in addition to the input node).

[0109] The artificial neural network also includes an output layer 307, which provides the output from the artificial neural network; that is, the output data of the artificial neural network is the node values ​​of the output layer. As for the hidden / processing layers, the output node values ​​are generated as a function of the node values ​​of the previous layers. However, unlike the hidden / processing layers, where the node values ​​are generally inaccessible or not further usable, the node values ​​of the output layer are accessible and provide the results of the artificial neural network's operations.

[0110] Several different network architectures and toolkits have been developed for artificial neural networks, and in many embodiments, artificial neural networks can be based on adapted and customized networks. An example of a network architecture suitable for the aforementioned applications is WaveNet by van den Oord et al., described in the following document: Oord, Aaron van den, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, “Wavenet: A Generative Model for Raw Audio.” arXiv preprint arXiv:1609.03499 (2016).

[0111] WaveNet is an architecture for synthesizing temporal signals using dilated causal convolutions, and it has been successfully applied to audio signals. For WaveNet, the following activation function is typically used: in Let represent the convolution operator, ⊙ represent the element-wise multiplication operator, σ(·) be the sigmoid function, k be the layer index, f and g represent the filter and gate, respectively, and W represent the weights of the learned artificial neural network. The filter product of an equation can often be used with a gate product to provide a filtering effect, thus providing a weighted result. This can effectively allow the contribution of a node to be reduced to essentially zero in many cases (i.e., it can allow or "prohibit" a node from contributing to other nodes, thus providing a "gate" function). In different cases, the gate function can result in the node's output being negligible, while in other cases it will make a significant contribution to the output. Such a function can substantially help allow the neural network to learn and be trained effectively.

[0112] In some cases, artificial neural networks can also be configured to include additional contributions that allow the network to be dynamically adapted or tailored to specific desired properties or characteristics of the generated output. For example, a set of values ​​can be provided to adapt the network. These values ​​can be included by contributing to some nodes of the network. These nodes can be specifically input nodes, but are typically nodes in hidden or processing layers. Such adaptation values ​​can be, for example, weighted and added as a contribution to a weighted sum / correlation value for a given node. For example, in WaveNet, such adaptation values ​​can be included in the activation function. For example, the output of the activation function can be given as: Where y is a vector representing the fit values, and V represents the appropriate weights for these values.

[0113] The above description relates to neural network methods applicable to many embodiments and implementations. However, it should be understood that many other types and structures of neural networks can be used. In fact, many different methods for generating neural networks have been and are being developed, including neural networks using complex structures and processes different from those described above. This method is not limited to any particular neural network method, and any suitable method can be used without departing from the invention.

[0114] The inventors have recognized that, by addressing specific problems / effects of many conventional methods, the specific structure and use of artificial neural networks to perform temporal envelope compensation provide a particular level of efficiency. The inventors have also recognized that generating decorrelated signals by means of, for example, all-pass processing is highly efficient for stationary signals, but for non-stationary signals, such as, for example, applause signals, the decorrelated signals often inappropriately follow the temporal envelope of the received downmixed audio signal. Figure 7 The image shows an example of a mono downmixed audio signal and the resulting decorrelation signal. It can be observed that most of the temporal detail or sharpness is lost in the decorrelation signal. This results in a decrease in audio quality after upmixing the downmixed audio signal to generate a multichannel audio signal based on the upmixing parameters.

[0115] A decorrelated signal that more closely follows the time envelope of the downmix / mono signal is preferred, and... Figure 1 In the device, this is achieved through envelope compensation, including envelope compensator 109.

[0116] If the time-domain representation of a mono downmixed audio signal is expressed as a function of the envelope... Modulated time-flat signal :

[0117] Correlation signal generated accordingly The time-domain representation is represented by the envelope. Modulated time-flat signal :

[0118] The desired decorrelation signal is close to the following: (3)

[0119] One solution could include: converting the decorrelated signal back to the time domain using a hybrid synthesis filter bank (and converting the downmixed audio signal if it is not yet available in the time domain), running two envelope estimation algorithms separately on the time-domain downmixed audio signal and the time-domain decorrelated signal, adjusting the time envelope of the decorrelated signal accordingly, and converting the adaptive decorrelated signal back to the hybrid QMF domain using a hybrid analysis group. However, this operation is computationally expensive and introduces significant processing latency.

[0120] To mitigate the problem, a solution operating in the frequency domain, and specifically in the mixed QMF domain, can be employed. Specifically, the envelope compensator 109 can use a pre-trained artificial neural network to recover the temporal envelope of the downmixed audio signal to the decorrelated signal. This method can specifically provide improved temporal envelope reconstruction with typically, and particularly, finer temporal resolution. This is achieved using an artificial neural network trained to learn how information from different frequency samples should be combined (in both temporal and frequency directions) to reconstruct a temporally high-resolution envelope.

[0121] In some embodiments, the envelope compensator 109 may include a single trained artificial neural network that directly generates samples of the compensated decorrelation signal. Therefore, the trained artificial neural network may have output nodes that output samples of the compensated decorrelation signal.

[0122] Typically, a trained artificial neural network generates frequency samples (specifically, complex frequency slot / frequency interval values) of the compensated decorrelation signal, and in many embodiments, the artificial neural network directly generates the compensated frequency-domain decorrelation signal. However, in some embodiments, the artificial neural network can generate time-domain samples of the compensated decorrelation signal. In such examples, these time-domain samples may be converted to the frequency domain before being provided to generator 105 for upmixing. In many embodiments, a trained artificial neural network can be directly configured to generate frequency samples of the compensated frequency-domain decorrelation signal and then fed to generator 105 for upmixing. Therefore, in many embodiments, the trained artificial neural network of envelope compensator 109 directly generates the frequency-domain decorrelation signal.

[0123] In many embodiments, the trained artificial neural network directly receives frequency samples of the frequency-domain decorrelation signal and the frequency-domain audio signal. Therefore, the trained artificial neural network can have input nodes that directly receive frequency samples of the frequency-domain decorrelation signal and the frequency-domain audio signal.

[0124] In many embodiments, the input nodes of the trained artificial neural network receive frequency-domain signal samples of the frequency-domain audio signal and the frequency-domain decorrelation signal, respectively, and the output nodes provide frequency samples of the compensated frequency-domain decorrelation signal. However, although the artificial neural network operates in the frequency domain by receiving frequency-domain input and generating frequency-domain output, the function of the trained artificial neural network is to reduce the temporal envelope between the input and output decorrelation signals relative to the downmixed audio signal. This is achieved by training the network to provide an output with a compensated temporal envelope. For example, as will be described in more detail later, the training of the artificial neural network can be based on a cost function that reflects the envelope difference between the temporal signal / representation of the compensated frequency-domain decorrelation signal and the temporal signal / representation of the frequency-domain audio signal for a large number of training examples.

[0125] exist Figure 1 In the envelope compensator 109, compensation for the time-domain envelope characteristics can therefore be achieved without any transformation to the time domain and when operating only on frequency-domain samples. This is achieved by using a trained artificial neural network that can not only be trained to reflect time-domain differences but also inherently incorporates characteristics of time-domain and frequency-domain transformations / relationships.

[0126] exist Figure 5 An example of this method is shown. In this example, the envelope compensator 109 and its trained artificial neural network do not explicitly generate intermediate time-domain envelopes or signals. Instead, in this example, the envelope compensator 109 essentially comprises a single trained artificial neural network, which includes a network architecture trained to process input frequency samples of a given frequency-domain audio signal. Frequency samples of frequency-domain decorrelated signals In this case, frequency samples of the compensated frequency domain decorrelation signal are directly provided. .

[0127] In some embodiments, the envelope compensator 109 may be arranged to generate a time-domain envelope estimate (signal) of at least one of a frequency-domain audio signal and a frequency-domain decorrelation signal, and then generate a sample of the compensated frequency-domain decorrelation signal based on the time-domain envelope estimate signal.

[0128] In such embodiments, one or both of the time-domain envelope estimation signal estimators may be implemented as trained artificial neural networks, and / or the function that determines samples of the compensated frequency-domain decorrelation signal based on the time-domain envelope estimation signal may be a trained artificial neural network.

[0129] Figure 6 An example of an envelope compensator 109 is shown, which includes a first envelope estimator 601 arranged to generate a first time-domain envelope estimate and a second envelope estimator 603 arranged to generate a second time-domain envelope estimate, wherein the time-domain envelope estimates depend on the properties of the time-domain envelopes of the frequency-domain audio signal and the frequency-domain decorrelation signal, respectively.

[0130] In this example, envelope compensator 109 also includes compensator 605, which is arranged to generate a sample of the compensated frequency domain decorrelation signal from the first and / or second time domain envelope estimates and the generally frequency domain decorrelation signal itself.

[0131] In many embodiments, the first time-domain envelope estimate is a signal representing a time-domain envelope estimate of a frequency-domain audio signal, and / or the second time-domain envelope estimate is a signal representing a time-domain envelope estimate of a frequency-domain decorrelation signal.

[0132] In some embodiments, the first and / or second time-domain envelope estimates may be provided in the form of features / parameters that depend on / reflect the properties or characteristics of the time-domain envelope of the frequency-domain audio signal and / or the frequency-domain decorrelated signal, respectively. For example, the time-domain envelope may be represented by features or parameters, including, for example, average level, intermediate level, maximum level, minimum level, a measure of change such as median or average frequency, etc.

[0133] In some embodiments, the first envelope estimator 601 is implemented as an artificial neural network, hereinafter referred to as the first artificial neural network, which receives samples of a frequency-domain audio signal as input and generates a time-domain envelope estimation signal of N samples / values ​​providing data representing the attributes of the time-domain envelope of the frequency-domain audio signal. The artificial neural network can be trained using training data comprising training frequency-domain audio signals with associated, determined time-domain envelopes, the cost function reflecting the difference between a metric of the time-domain envelope of the resulting compensated frequency-domain decorrelation signal and a corresponding metric of the time-domain envelope of the frequency-domain audio signal. Therefore, the cost function can include a comparison of the time-domain envelopes of the frequency-domain audio signal and the compensated frequency-domain decorrelation signal.

[0134] Alternatively or additionally, in some embodiments, the second envelope estimator 603 may be implemented as an artificial neural network, hereinafter referred to as the second artificial neural network, which receives samples of the frequency-domain decorrelated signal as input and generates a time-domain envelope estimation signal of N samples / values ​​providing attributes of the time-domain envelope of the frequency-domain decorrelated signal. The artificial neural network can be trained using training data comprising the training frequency-domain decorrelated signal and a frequency-domain audio signal having a defined time-domain envelope associated with the frequency-domain audio signal, the cost function reflecting the difference between a metric of the time-domain envelope of the resulting compensated frequency-domain decorrelated signal and a corresponding metric of the time-domain envelope of the frequency-domain audio signal. Therefore, the cost function may include a comparison of the time-domain envelopes of the frequency-domain audio signal and the compensated frequency-domain decorrelated signal.

[0135] In examples where the first envelope estimator 601 and / or the second envelope estimator 603 are implemented as artificial neural networks, the temporal envelope estimation can, for example, be generated as samples of the time-domain or frequency-domain signal corresponding to the temporal envelope. In other embodiments, the temporal envelope estimation can alternatively or additionally be generated as a set of parameters. Such parameters can directly correspond to specific properties of the temporal envelope, such as frequency, average level, etc. However, in other embodiments, the generated features can be more abstract and not directly relate to specific features. Instead, the first envelope estimator 601 and / or the second envelope estimator 603 can be artificial neural networks trained to generate features that, for example, allow the compensator 605 (e.g., implemented as an artificial neural network) to generate improved compensated frequency decorrelation signals. In this case, the parameters / features generated by the first envelope estimator 601 and / or the second envelope estimator 603 can represent abstract or unknown properties that allow the artificial neural network to provide efficient performance (based on training).

[0136] In other embodiments, the first envelope estimator 601 and / or the second envelope estimator 603 may be implemented for more conventional, predetermined, and / or analytical functions. For example, the first envelope estimator 601 or the second envelope estimator 603 may be implemented as a frequency-to-time domain transformer followed by a low-pass filter with a low cutoff frequency to generate a time-domain envelope signal. These can then be used as inputs to a trained artificial neural network that forms the compensator 605.

[0137] It should also be understood that the first envelope estimator 601 or the second envelope estimator 603 may or may not be implemented as the same or similar circuitry. For example, in some embodiments, one of the envelope estimators may be implemented by a trained artificial neural network, while the other may be implemented as a conventional envelope estimator. In other embodiments, both the first envelope estimator 601 and the second envelope estimator 603 may be implemented as artificial neural networks, but with different properties, such as, for example, different numbers of input or output nodes, different numbers of layers, different numbers of nodes in each layer, etc.

[0138] In some embodiments, the compensator 605 may be implemented as an artificial neural network, hereinafter referred to as a compensator artificial neural network, which receives samples from the first envelope estimator 601 and the second envelope estimator 603, and continues to generate samples of the compensated frequency-domain decorrelation signal. Typically, the compensator artificial neural network also receives samples of the frequency-domain decorrelation signal. Thus, the compensator artificial neural network may have input nodes for receiving samples of the frequency-domain decorrelation signal and the outputs of the first envelope estimator 601 and the second envelope estimator 603 (such as, specifically, the time-domain envelope estimate of the frequency-domain decorrelation signal and the frequency-domain audio signal).

[0139] Similar to the first envelope estimator 601 and the second envelope estimator 603 artificial neural networks, the compensator artificial neural network can be trained by a large number of training frequency-domain audio signals (and frequency-domain decorrelation signals generated therefrom) as inputs to the envelope compensator 109, and the compensator artificial neural network can be trained using a cost function that compares the time-domain envelope of the resulting compensated frequency-domain decorrelation signal with the time-domain envelope of the frequency-domain audio signal.

[0140] It should be understood that the envelope compensator 109 may use a trained artificial neural network to implement only one or two of the functions of the first envelope estimator 601, the second envelope estimator 603, and the compensator 605. In such cases, the given trained artificial neural network may be trained independently while using the expected or nominal implementation of the other functions.

[0141] A particular advantageous method can be implemented by using a first artificial neural network as a first envelope estimator 601, a second artificial neural network as a second envelope estimator 603, and a third artificial neural network as a compensator 605. This method has been found to provide particularly accurate and fine-grained resolution envelope compensation while ensuring reduced computational complexity and resource usage. In particular, the combined complexity of the three individual artificial neural networks is generally much less than the complexity required to achieve similar accuracy using a single artificial neural network attempting to perform full temporal envelope compensation. The use of artificial neural networks in a specific architecture with specific interactions and intermediate temporal envelope estimations for specific purposes has been found to be highly effective and provides accurate envelope compensation.

[0142] In many embodiments, the envelope compensator 109 may include a neural network architecture comprising two temporal envelope estimators and an envelope recovery / compensation process, all implemented by a neural network.

[0143] As a concrete example, for frequency domain audio signals and frequency domain decorrelation signals (e.g., in the mixed QMF signal domain), an artificial neural network can be pre-trained to make it suitable for complex-valued mixed QMF samples ( Each frequency band and A block of (number of time slots) is predicted by an artificial neural network using two temporal envelope estimation techniques. The corresponding temporal envelope of each sample. In some embodiments, the architecture and the trained coefficients of the artificial neural network that is actually the temporal envelope estimator can be the same or can be shared. However, in other cases, there may be differences, and the artificial neural network can actually be trained separately.

[0144] The compensator artificial neural network is trained to generate samples of compensated frequency domain decorrelation signals, specifically generating hybrid QMF samples. Decorrelated mixed QMF blocks And the generated temporal envelope. In Figure 6 In the example, a single time slot is generated ( However, artificial neural networks can also be configured to generate multiple time slots. Artificial neural networks can be trained to generate processed decorrelational mixed QMF samples such that, in the time domain, the resulting envelope-compensated signal resembles: .

[0145] A possible network architecture is the so-called fully connected (FC) network architecture, such as... Figure 8 As shown. In this example, in the first step, blocks of complex-valued mixed subband samples are flattened to form vectors of (real-valued) mixed subband domain samples. Then, in each FC layer, a set of weights and nonlinear elements transform the input vector into a vector in the next layer, until the final output generates the desired output vector of samples of the compensated frequency-domain decorrelation signal.

[0146] The processing of the envelope compensator 109, and indeed most or all of the process of generating the multichannel audio signal, is typically performed as segmented processing, where blocks / time intervals of samples of the multichannel audio signal are processed. In particular, a trained artificial neural network can perform block processing, where a set / block of output samples is generated based on a set / block of input values ​​generated from a set / block of samples of the downmixed audio signal and the decorrelated signal.

[0147] In some embodiments, the time interval for determining the compensated frequency-domain decorrelation signal for a sample within an operation / block / processing segment is the same as the time interval for the frequency-domain audio signal / frequency-domain decorrelation signal from which the input value is generated; that is, in some embodiments, the time interval for the output value matches the time interval for the input value. However, in other embodiments, the output value for a given time interval may be determined based on or from the input value of the input signal (frequency-domain audio signal and frequency-domain decorrelation signal) with a larger time interval, typically a time interval whose duration is not less than 2, 3, 5, or 10 times the time interval for generating the sample.

[0148] Specifically, when generating a sample block of compensated frequency-domain decorrelation signal in a processing segment / operation, the processing and input to the artificial neural network can be based on input values ​​from multiple processing time intervals. For example, when determining the samples of the compensated frequency-domain decorrelation signal for the current processing segment, the input values ​​can include not only sample values ​​of the frequency-domain audio signal and frequency-domain decorrelation signal for the current processing segment's time interval, but also sample values ​​from previous and, for example, the next processing time interval. As an example, in Figure 6In the method, when generating a compensated frequency domain decorrelation signal for one processing segment / block, the inputs of the artificial neural networks of the first envelope estimator 601 and the second envelope estimator 603 include S frequency samples of the processing segment / block.

[0149] In many embodiments, the envelope compensator 109 may be arranged such that samples of the frequency-domain audio signal and / or the frequency-domain decorrelation signal are directly fed to the input nodes of the appropriate artificial neural network. However, in some embodiments, at least some of the input values ​​of the artificial neural network may not be directly signal samples, but may be derived from signal samples. Specifically, some or all of the input nodes may instead (or possibly also) receive individual sample values ​​having characteristics of the signal, wherein the characteristics are determined from the signal samples.

[0150] This feature determination can be performed through dedicated analysis or predetermined functions, such as the absolute value of the subband signal or the first-order difference between subband samples, to highlight transient components in the frequency domain audio signal. Predetermined functions may include, for example, transient detectors, envelope trackers, stationarity indicators that use different smoothing filters to detect envelope changes, etc.

[0151] In many embodiments, feature determination itself can be performed by an artificial neural network architecture, or can actually be implemented as part of the artificial neural network itself. For example, the artificial neural network may include initial 1D or 2D convolutions to preprocess the input signal through multiple feature maps before applying the input signal to the FC layer.

[0152] In other embodiments, the features determined by 1D or 2D convolution can themselves be frequency domain signals, where different properties of the frequency domain (subband) waveform are emphasized or de-emphasized, depending on what the network has learned is necessary to properly perform envelope compensation. The emphasis / de-emphasis of waveform components such as those with abrupt energy changes can also depend on the inter-frequency correlation of these components (regarding their importance in determining the required signal transformation to produce the desired time-domain envelope compensation).

[0153] In some embodiments, the envelope compensator 109 may be specifically arranged to perform envelope compensation in the feature domain rather than the signal domain. Figure 10 An example of this method is shown in the figure.

[0154] In this example, the envelope compensator 109 includes a first feature determiner 901, which is arranged to determine a first set of features from a frequency domain decorrelation signal. The envelope compensator 109 further includes a second feature determiner 903, which is arranged to determine a second set of features from a frequency domain audio signal.

[0155] The first feature determiner 901 and / or the second feature determiner 903 can be implemented as a trained artificial neural network. In other embodiments, they can be implemented as low-complexity or simplified artificial neural networks, such as one or more convolutional layers.

[0156] Specifically, in a particular example, the frequency domain audio signal M[k, m] and the frequency domain decorrelation signal D[km] can be obtained by extracting features from each signal through 1D / 2D convolutional layers.

[0157] In other embodiments, features may be determined, for example, using a predetermined or analytical function.

[0158] Figure 9 The envelope compensator 109 also includes a trained artificial neural network, hereinafter referred to as the adjustment estimation neural network 905, which is arranged to generate a set of feature scaling factors based on the determined features. The adjustment estimation neural network has input nodes that receive a first feature set and a second feature set from a first feature determiner 901 and a second feature determiner 903. It also has output nodes that provide the set of feature scaling factors.

[0159] Figure 9 The envelope compensator 109 also includes an adjustment circuit 907 arranged to perform an adjustment of a first feature set, and in this example, specifically applies a scaling factor to the first feature set to generate a first adjusted / compensated feature set. Specifically, each feature (or at least one or more features) in the first feature set representing the frequency-domain decorrelated signal is scaled based on a scaling factor determined by the adjustment estimation neural network. Therefore, the adjustment / scaling is performed in the feature domain, and it is scaling of the features rather than scaling of the signal samples.

[0160] In a specific example, the features are scaled based on a generated scaling factor and then summed with the original values; that is, the total applied scaling is 1 + G, where G is the gain value determined by the adjusted estimation neural network for the given features. Alternatively, the adjustment circuit 907 can consist of a simple nonlinear layer, where the output is a nonlinearly scaled weighted sum of the inputs.

[0161] The scaled feature values ​​are then fed into a feature decoder 909, which is configured to generate an envelope-compensated frequency-domain decorrelation signal from the first compensated feature set. Therefore, after scaling / compensation, features representing the frequency-domain decorrelation signal are fed into the feature decoder 909, which continues to generate output samples of the compensated frequency-domain decorrelation signal.

[0162] In many embodiments, the feature decoder 909 can be implemented as a trained artificial neural network. In particular, it has been found that using a properly trained artificial neural network can achieve efficient construction of the compensated frequency domain decorrelation signal based on the features.

[0163] This method provides efficient and accurate envelope compensation in many embodiments and for many signals. The architecture is particularly effective for implementations using artificial neural networks, and in particular, it has been found that compensation in the feature domain is performed based on the evaluation of features of both the frequency-domain decorrelation signal and the frequency-domain audio signal by the artificial neural network, allowing for highly efficient implementation and high performance.

[0164] To further clarify Figure 9 The method of envelope compensator 109, as previously described, can be considered as follows:

[0165] The recovered envelope decorrelation signal can be written based on the original term as:

[0166] If Divide by If the quotient is 1, then the remainder is And output:

[0167] Use learning gain or mask term replace And move it to the feature domain to get:

[0168] In some embodiments, additional processing may be applied to the gain ratio. This prevents large or near-zero scaling factors. For example, a non-linear function can be applied to keep the ratio within a given range.

[0169] exist Figure 9 In the neural network architecture, the mask term The feature can be determined by the operations of the first feature determiner 901, the second feature determiner 903, and the adjusted estimation neural network 905. This can then be applied to the representation. The first set of features is then added to the same input features. The features can then be decoded via the deconvolution layer of the feature decoder 909 to produce... Therefore, compared to directly learning scaling... In comparison, item It can be regarded as a correction factor and is generally easier to model using neural networks.

[0170] It should be understood that while the example specifically determines a scaling factor and adjusts the features by scaling, other adjustments to the features are possible. For example, in some embodiments, an adjustment estimation neural network can be trained to determine the offset of the features, and the determined offset can be added to the adjustment circuitry that performs summation instead of scaling. It will also be understood that such variations or alternatives can be achieved by adapting the training process by including the adjustment circuitry operations in the training of the adjustment estimation neural network.

[0171] In some embodiments, envelope compensation can be performed across the entire frequency range, and specifically, compensation can be applied to all frequency sub-bands. However, in other embodiments, the envelope compensator 109 can be arranged to generate an envelope-compensated frequency-domain decorrelation signal by applying time-domain envelope compensation only to a subset of the higher frequency sub-bands of the frequency-domain decorrelation signal. Therefore, specifically, in some cases, envelope compensation can be limited to higher frequency sub-bands, and specifically limited to frequency sub-bands above a given threshold frequency.

[0172] This method can be particularly well combined with hybrid frequency representations, where subbands can have different bandwidths. For example, as... Figure 2 As shown, the time-domain to frequency-domain transformation can be applied to generate a common frequency transform that produces a set of subbands with equal subbands. These lower-frequency subbands can then be further subdivided to generate subbands with lower bandwidths. A delay can be introduced into the higher-frequency subbands (to compensate for the processing time of further subdividing the lower-frequency subbands). In such an example, envelope compensation can be applied to one or more higher-frequency bands but not to any further subdivided lower-frequency subbands. In this case, envelope compensation can be performed in parallel with the delay of the higher-frequency bands. Therefore, envelope compensation can be performed on the higher-frequency bands / signal before the delay. In many cases, the envelope compensation operation can be faster than frequency subdivision, and therefore, this method can allow envelope compensation to be performed without introducing any additional delay.

[0173] Figure 10 It shows having Figure 2 Time-domain to frequency conversion function Figure 1 An example of this method in an audio device.

[0174] This method leverages the understanding that the fine details of the time-domain / time-domain envelope are largely determined by higher frequencies, and therefore compensation can be focused on higher frequencies. This allows for, for example... Figure 10 The structure includes envelope compensation without introducing additional delay in the signal path. In this example, the lower frequency band includes additional frequency subdivision but no envelope compensation. Decorrelation is applied to the higher and lower frequency paths, respectively.

[0175] For higher frequency bands, the decorreductor 107 is fed the signal before delay compensation of the hybrid filter bank. Similarly, the envelope compensator 109 is executed in parallel with the delay of the signal path, and the envelope compensation can be arranged so as not to add any additional delay to the signal path, provided that the processing time required for decorreduction and envelope compensation is less than the delay of the additional frequency subdivision.

[0176] It should be understood that many variations are possible. The low-frequency band of the frequency domain decorrelation signal can be compensated using control data determined from the high-frequency band, thereby reducing the delay of the signal path used for the lower frequency band. In some embodiments, the processing for, for example, a frequency band can be asymmetric. For example, the higher frequency band can use both look-ahead and look-back buffers to estimate samples, while the lower frequency band can use only the look-back buffer.

[0177] Artificial neural networks are adapted for specific purposes through training processes that adapt / tune / modify the weights and other parameters (e.g., biases) of the artificial neural network. It should be understood that many different training processes and algorithms are known for training artificial neural networks. Figure 11 The diagram shows a training setup that can be used for the artificial neural network described.

[0178] Typically, training is based on a large training set, which provides the network with numerous examples of input data. Furthermore, the output of an artificial neural network is usually (directly or indirectly) compared to an expected or desired result. A cost function can be generated to reflect the desired outcome of the training process. In a typical scenario known as supervised learning, the cost (or loss) function typically represents the distance between the prediction and the ground truth for a given input data. Based on the cost function, the weights can be changed, and by repeating this process with the modified weights, the artificial neural network can be fitted to a state that minimizes the cost function.

[0179] Different methods can be used to train the neural network, and specifically, overall training can seek to result in the output of the audio device being the most closely corresponding multichannel audio signal to the original multichannel audio signal. Therefore, the arrangement can be trained to provide a compensated frequency-domain decorrelation signal that most effectively leads to an accurate reconstruction of the multichannel audio signal. In this case, a cost function based on the difference between the generated multichannel audio signal and the original training multichannel audio signal can be used. In other cases, more specific training of the artificial neural network can be used. For example, the training of the artificial neural network of envelope compensator 109 can be based on a cost function reflecting the difference between the time-domain envelope of the generated compensated frequency decorrelation signal and the time-domain envelope of the corresponding frequency-domain audio signal. In other cases, even more specific training can be used, where, for example, the artificial neural network of the first envelope estimator 601 is performed based on a cost function comparing the time-domain envelope of the training frequency-domain audio signal and the estimated time-domain envelope of the generated signal. In some embodiments, general and joint training of multiple artificial neural networks can be performed (e.g., using a cost function that evaluates the generated multichannel audio signal), and in other embodiments, individual artificial neural networks can be trained separately.

[0180] In some embodiments, the cost function may, for example, compare the training target frequency domain signal (based on a pre-computed time domain signal with a corrected / matched envelope) with the frequency domain signal generated by the envelope compensator 109.

[0181] More specifically, during the training step, a neural network can have two distinct information flows: from input to output (forward propagation) and from output to input (backward propagation). In the forward propagation, as described above, the neural network processes the data, while in the back propagation, the weights are updated to minimize the cost function. Typically, this backpropagation follows the gradient direction of the cost function landscape. In other words, by comparing the predicted output with the ground truth of a batch of data inputs, the direction in which the cost function is minimized and backpropagation proceeds can be estimated by updating the weights accordingly. Other known methods for training artificial neural networks include, for example, the Levenberg-Marquardt algorithm, the conjugate gradient method, and Newton's method.

[0182] In the current context, training may specifically include a training set comprising a potentially large number of multichannel audio signals or corresponding downmixed audio signals. The training set may include audio signals representing multiple different audio sources, including, for example, video recordings, movies, telecommunications, etc. In some embodiments, the training data may even include non-audio data, such as training performed in combination with training data from other sources (such as text data, etc.).

[0183] In some embodiments, the training data may be a multichannel audio signal within a time period corresponding to the processing time interval of the artificial neural network being trained. For example, the number of samples in the training multichannel audio signal may correspond to the number of samples corresponding to the input nodes of the artificial neural network being trained. Thus, each training example may correspond to one operation of the artificial neural network being trained. However, typically, a batch of training samples is considered for each step to accelerate and smooth the training process. Furthermore, numerous upgrades to gradient descent are also possible to accelerate convergence or avoid local minima in the overview of the cost function.

[0184] For each training multichannel audio signal, the training processor can perform a downmixing operation to generate a downmixed audio signal and corresponding upmixing parameter data. Therefore, the encoding process applied to multichannel audio signals during normal operation can also be applied to training multichannel audio signals to generate downmixing and upmixing parameter data.

[0185] Furthermore, in some embodiments, the training processor can generate a time-domain envelope signal for the frequency-domain audio signal.

[0186] The training downmixed audio signal can be fed into an arrangement of audio devices including a decorrelator 107 and an envelope compensator 109. The output of the neural network operation from the arrangement is then determined, and a cost function is applied to determine the cost value of each training downmixed audio signal and / or a set of combined training downmixed audio signals (e.g., determining the average cost value of the training set). The cost function can include various components.

[0187] Typically, the cost function will include at least one component reflecting how close the generated signal is to the reference signal, i.e., the so-called reconstruction error. In some embodiments, the cost function will include at least one component reflecting how close the generated signal is to the reference signal from a perceptual perspective.

[0188] For example, in some embodiments, the time-domain envelope of the compensated frequency-decorrelated signal generated by the envelope compensator 109 for a given training mixed audio signal / multichannel audio signal can be compared with the time-domain envelope of the training mixed audio signal. The time-domain envelope estimated signal can be determined by transforming it to the time domain of a low-pass filtered version of the signal. The resulting signals can be compared to determine the cost function. This process can be generated for all training sets to generate the total cost function.

[0189] Based on cost values, a training processor can adapt the weights of an artificial neural network. For example, backpropagation can be used. Specifically, the training processor can adjust the weights of one or all of the artificial neural networks based on cost values. For instance, given the derivative of the weights with respect to the cost function (representing the slope), the weight values ​​are modified along the slope direction. For a simple / minimum interpretation, refer to training a perceptron (a single neuron) using backpropagation on a single data input.

[0190] This process can be iterated until the artificial neural network is considered trained. For example, training can be performed for a predetermined number of iterations. As another example, training can continue until the weights change by less than a predetermined amount. Also very common is to implement validation stopping, where the network validation metric is tested again and stops when the expected result is reached.

[0191] In some embodiments, the cost function may be arranged to reflect the difference between the training multichannel audio signal and the multichannel audio signal generated by upmixing the downmixed audio signal with a compensated frequency decorrelation signal. As previously mentioned, in other embodiments, the cost function may reflect differences in intermediate parameters or signals, such as the difference in the time-domain envelope between the frequency-domain audio signal and the compensated frequency decorrelation signal. In some cases, the cost function may include contributions from multiple such comparisons.

[0192] More specifically, the cost function may include the (root) mean square error between the frequency domain output signal of the artificial neural network and the training target time domain decorrelation signal transformed in the frequency domain, wherein the time domain envelope of the training signal matches the time domain envelope of the audio signal.

[0193] Audio devices can be specifically implemented in one or more appropriately programmed processors. In particular, artificial neural networks can be implemented in one or more such appropriately programmed processors. Different functional blocks (especially artificial neural networks) can be implemented in separate processors and / or, for example, in the same processor. Examples of suitable processors are provided below.

[0194] Figure 12 This is a block diagram illustrating an example processor 1200 according to an embodiment of the present disclosure. Processor 1200 can be used to implement one or more processors, which implement the apparatus or elements thereof as described above (particularly including one or more artificial neural networks). Processor 1200 can be any suitable processor type, including but not limited to microprocessors, microcontrollers, digital signal processors (DSPs), field-programmable arrays (FPGAs) (wherein a field-programmable gate array has been programmed to form a processor), graphics processing units (GPUs), application-specific integrated circuits (ASICs) (wherein an ASIC has been designed to form a processor), or combinations thereof.

[0195] Processor 1200 may include one or more cores 1202. Core 1202 may include one or more arithmetic logic units (ALUs) 1204. In some embodiments, in addition to or in place of ALU 1204, core 1202 may include a floating-point logic unit (FPLU) 1206 and / or a digital signal processing unit (DSPU) 1208.

[0196] Processor 1200 may include one or more registers 1212 communicatively coupled to core 1202. Registers 1212 may be implemented using dedicated logic gates (e.g., flip-flops) and / or any memory technology. In some embodiments, registers 1212 may be implemented using static memory. Registers may provide data, instructions, and addresses to core 1202.

[0197] In some embodiments, processor 1200 may include a level-one or multi-level cache memory 1210 communicatively coupled to core 1202. Cache memory 1210 may provide computer-readable instructions to core 1202 for execution. Cache memory 1210 may provide data for core 1202 to process. In some embodiments, computer-readable instructions may have already been provided to cache memory 1210 by local memory (e.g., local memory attached to external bus 1216). Cache memory 1210 may be implemented using any suitable cache memory type, such as metal-oxide-semiconductor (MOS) memory, such as static random-access memory (SRAM), dynamic random-access memory (DRAM), and / or any other suitable memory technology.

[0198] Processor 1200 may include controller 1214, which can control inputs to processor 1200 from other processors and / or components included in the system and / or outputs from processor 1200 to other processors and / or components included in the system. Controller 1214 can control data paths in ALU 1204, FPLU 1206, and / or DSPU 1208. Controller 1214 may be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of controller 1214 may be implemented as stand-alone gates, FPGAs, ASICs, or any other suitable technology.

[0199] Register 1212 and cache 1210 can communicate with controller 1214 and core 1202 via internal connections 1220A, 1220B, 1220C and 1220D. The internal connections can be implemented as buses, multiplexers, cross switches and / or any other suitable connection technology.

[0200] The processor 1200's inputs and outputs may be provided via bus 1216, which may include one or more wires. Bus 1216 may be communicatively coupled to one or more components of the processor 1200, such as controller 1214, cache 1210, and / or register 1212. Bus 1216 may be coupled to one or more components of the system.

[0201] Bus 1216 may be coupled to one or more external memories. The external memory may include read-only memory (ROM) 1232. ROM 1232 may be a mask ROM, electrically programmable read-only memory (EPROM), or any other suitable technology. The external memory may include random access memory (RAM) 1233. RAM 1233 may be static RAM, backup battery static RAM, dynamic RAM (DRAM), or any other suitable technology. The external memory may include electrically erasable programmable read-only memory (EEPROM) 1235. The external memory may include flash memory 1234. The external memory may include a magnetic storage device, such as a disk 1236. In some embodiments, the external memory may be included in the system.

[0202] This invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. The invention can optionally be implemented, at least in part, as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the invention can be implemented physically, functionally, and logically in any suitable manner. In practice, functionality can be implemented in a single unit, in multiple units, or as part of other functional units. Therefore, the invention can be implemented in a single unit or can be physically and functionally distributed among different units, circuits, and processors.

[0203] Although the invention has been described in conjunction with some embodiments, it is not intended to be limited to the specific forms set forth herein. Rather, the scope of the invention is limited only by the appended claims. Furthermore, while features may appear to have been described in conjunction with specific embodiments, those skilled in the art will recognize that various features of the described embodiments can be combined according to the invention. In the claims, terms include those that do not exclude the presence of other elements or steps.

[0204] Furthermore, although listed separately, multiple units, elements, circuits, or method steps can be implemented by, for example, a single circuit, unit, or processor. Additionally, although individual features may be included in different claims, these features may be advantageously combined, and inclusion in different claims does not imply that such combinations are impractical and / or advantageous. Moreover, including a feature in one class of claims does not imply a limitation on that class, but rather indicates that the feature is equally applicable to other claim classes. Furthermore, the order of features in a claim does not imply that these features must operate in any particular order, and in particular, the order of steps in a method claim does not imply that the steps must be performed in that order. Rather, the steps can be performed in any suitable order. Furthermore, singular references do not exclude plural. Therefore, references to “a,” “an,” “first,” “second,” etc., do not exclude plural. Reference numerals in the claims are provided as clarifying examples only and should not be construed as limiting the scope of the claims in any way.

Claims

1. An apparatus for generating multi-channel audio signals, the apparatus comprising: A receiver (101) is configured to receive frequency domain audio signals; A decorrelation unit (107) is arranged to generate a frequency domain decorrelation signal by applying decorrelation to the frequency domain audio signal; An envelope compensator (109) is arranged to generate an envelope-compensated frequency-domain decorrelation signal by applying time-domain envelope compensation to the frequency-domain decorrelation signal, the time-domain envelope compensation reducing the difference between the time-domain envelope of the frequency-domain audio signal and the time-domain envelope of the frequency-domain decorrelation signal. The envelope compensator (109) includes at least one trained artificial neural network having input values ​​determined based on at least one of the frequency-domain audio signal and the frequency-domain decorrelation signal and generating time-domain envelope compensation data for the time-domain envelope compensation, the envelope-compensated frequency-domain decorrelation signal depending on the time-domain envelope compensation data. as well as An output generator (105) is arranged to generate the multi-channel audio signal by upmixing the frequency domain audio signal and the envelope-compensated frequency domain decorrelated signal.

2. The apparatus according to claim 1, wherein, The at least one trained artificial neural network includes a trained artificial neural network arranged to receive an input value determined based on the frequency domain audio signal and to generate time-domain envelope compensation data indicating a time-domain envelope estimate of the frequency domain audio signal. The envelope compensator (109) is arranged to generate the envelope-compensated frequency-domain decorrelation signal based on the time-domain envelope estimate of the frequency domain audio signal.

3. The apparatus according to claim 1 or 2, wherein, The at least one trained artificial neural network includes a trained artificial neural network arranged to receive an input value determined based on the frequency domain decorrelation signal and to generate time-domain envelope compensation data indicating a time-domain envelope estimate of the frequency domain decorrelation signal, wherein the envelope compensator (109) is arranged to generate the envelope-compensated frequency domain decorrelation signal based on the time-domain envelope estimate of the frequency domain decorrelation signal.

4. The apparatus according to any one of the preceding claims, wherein, The envelope compensator includes: A first envelope estimator (601) is configured to determine a time-domain envelope estimate of the frequency-domain audio signal; A second envelope estimator (603) is configured to determine a time-domain envelope estimate of the frequency-domain decorrelated signal; and The at least one trained artificial neural network includes a trained artificial neural network (605) configured to receive an input value determined based on the time-domain envelope estimation of the frequency-domain audio signal and an input value determined based on the time-domain envelope estimation of the frequency-domain decorrelation signal, and to generate time-domain envelope-compensated data including samples of the envelope-compensated frequency-domain decorrelation signal.

5. The apparatus according to any one of the preceding claims, wherein, At least one of the at least trained artificial neural networks is configured to receive an input value determined based on the frequency domain audio signal and an input value determined based on the frequency domain decorrelation signal, and to generate time-domain envelope-compensated data to include samples of the envelope-compensated frequency domain decorrelation signal.

6. The apparatus according to any one of the preceding claims, wherein, The at least one trained artificial neural network is configured to generate temporal envelope compensation data for a first time interval based on input values ​​determined according to samples of at least one of the frequency-domain audio signal and the frequency-domain decorrelation signal in a second time interval, the second time interval being different from the first time interval.

7. The apparatus according to any one of the preceding claims, wherein, The envelope compensator (109) includes: A first feature determiner (901) is arranged to determine a first feature set based on the frequency domain decorrelation signal; A second feature determiner (903) is arranged to determine a second feature set based on the frequency domain audio signal; Wherein, the at least one trained artificial neural network includes a trained artificial neural network arranged to generate a set of feature adjustment factors, the artificial neural network having an input including the first feature set and the second feature set; Furthermore, the envelope compensator (109) also includes: Circuit (907), which is arranged to apply the adjustment factor to the first feature set to generate a first adjusted feature set; and A feature decoder (909) is arranged to generate the envelope-compensated frequency-domain decorrelation signal from the first adjusted feature set.

8. The apparatus according to claim 7, wherein, The feature decoder (909) includes a trained artificial neural network having an input node that receives the first adjusted feature set and an output node that provides samples of the envelope-compensated frequency-domain decorrelation signal.

9. The apparatus according to claim 7 or 8, wherein, The first feature determiner (901) includes an artificial neural network having an input node that receives samples of the frequency domain decorrelation signal and an output node that provides the first feature set.

10. The apparatus according to any one of claims 7 to 9, wherein, The second feature determiner (903) includes an artificial neural network having an input node that receives samples of the frequency domain audio signal and an output node that provides the second feature set.

11. The apparatus according to any of the preceding claims, wherein, The at least one artificial neural network is trained using training data comprising a number of training frequency-domain audio signals and a cost function, the cost function depending on a measure of the difference between the time-domain envelope of the training frequency-domain audio signals and the time-domain envelope of a compensated frequency-domain decorrelation signal generated for the training frequency-domain audio signals.

12. The apparatus according to any one of the preceding claims, wherein, The at least one trained artificial neural network includes an input convolutional layer that receives samples of the frequency domain audio signal and the frequency domain decorrelation signal, and at least one fully connected hidden layer.

13. The apparatus according to any of the preceding claims, wherein, The envelope compensator (109) is arranged to generate the envelope-compensated frequency-domain decorrelation signal by applying the time-domain envelope compensation only to a subset of the higher-frequency subbands of the frequency-domain decorrelation signal.

14. A method for generating a multi-channel audio signal, the method comprising: Receive frequency domain audio signals; A frequency-domain decorrelated signal is generated by applying decorrelation to the frequency-domain audio signal. An envelope-compensated frequency-domain decorrelation signal is generated by applying time-domain envelope compensation to the frequency-domain decorrelation signal. The time-domain envelope compensation reduces the difference between the time-domain envelope of the frequency-domain audio signal and the time-domain envelope of the frequency-domain decorrelation signal. The time-domain envelope compensation includes at least one trained artificial neural network having input values ​​determined based on at least one of the frequency-domain audio signal and the frequency-domain decorrelation signal and generating time-domain envelope compensation data for the time-domain envelope compensation. The envelope-compensated frequency-domain decorrelation signal depends on the time-domain envelope compensation data. as well as The multi-channel audio signal is generated by upmixing the frequency domain audio signal and the envelope-compensated frequency domain decorrelated signal.

15. A computer program product comprising computer program code units, wherein when the program is run on a computer, the computer program code units are adapted to perform all the steps of claim 14.