Generation of multi-channel audio signal and audio data signal representing multi-channel audio signal

By using an artificial neural network generator to upmix the downmixed audio signal, the problems of distortion and high complexity in the reconstruction of multi-channel audio signals in the prior art are solved, and more efficient audio quality and resource utilization are achieved.

CN121925702APending Publication Date: 2026-04-24KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
KONINKLIJKE PHILIPS NV
Filing Date
2024-09-12
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing parametric stereo coding methods suffer from problems such as distortion, high complexity, and high data rate during multi-channel audio signal reconstruction, leading to decreased audio quality and inefficient resource utilization.

Method used

An artificial neural network generator is used to upmix the downmixed audio signal using upmixing parameters and neural network control data to generate a multi-channel audio signal. The neural network is then optimized using training data to improve audio quality and reduce complexity.

Benefits of technology

It improves the reconstruction quality and flexibility of multi-channel audio signals, reduces data rate and computational load, and provides a more efficient audio processing method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121925702A_ABST
    Figure CN121925702A_ABST
Patent Text Reader

Abstract

A decoder audio device comprises a receiver (101) and an audio signal generator (103). The receiver receives an audio data signal, the audio data signal comprising: a downmix audio signal that is a downmix of a first multi-channel audio signal; an upmix parameter set for a time-frequency segment of the downmix audio signal, the upmix parameter set comprising a level difference parameter, a correlation parameter and a phase difference parameter; and at least one transient parameter indicative of a transient characteristic of the first multi-channel audio signal. The audio signal generator generates an output multi-channel audio signal by upmixing the downmix audio signal according to the upmix parameter and the transient parameter. The decoder audio device uses an artificial neural network (105) having an input node that receives the upmix parameter and the transient parameter. The transient parameter has a different time-frequency resolution than the upmix parameter, and typically has a much roughened frequency resolution. An improved multi-channel audio signal can be generated with little need for additional overhead in terms of processing complexity, resources, and data rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the generation of multi-channel audio signals and / or audio data signals representing multi-channel audio signals, and particularly, but not exclusively, to the encoding and / or decoding of stereo signals. Background Technology

[0002] Spatial audio applications have become numerous and widespread, increasingly becoming at least part of many audiovisual experiences. In fact, the continuous development of new and improved spatial experiences and applications has led to an increasing demand for audio processing and rendering.

[0003] For example, virtual reality (VR) and augmented reality (AR) have received increasing attention in recent years, and many implementations and applications are entering the consumer market. In fact, devices are being developed for rendering experiences and capturing or recording suitable data for these applications. For example, relatively low-cost devices are being developed to allow game consoles to provide a complete VR experience. With the VR and AR markets reaching a considerable size in a short period, this trend is expected to continue and, in fact, accelerate. In the audio domain, a prominent area of ​​exploration is the reproduction and synthesis of realistic and natural spatial audio. The ideal goal is to generate natural audio sources so that users cannot distinguish between synthesized and original audio sources.

[0004] Much research and development has focused on providing efficient and high-quality audio coding and decoding for spatial audio. Common spatial audio representations are multi-channel audio representations, including stereo representations, and efficient coding techniques for such multi-channel audio have been developed based on downmixing the multi-channel audio signal into a downmixed channel with fewer channels. One of the major advances in low-bit-rate audio coding is the use of parametric multichannel coding, where the downmixed signal is generated along with parametric data, which can be used to upmix the downmixed signal to reconstruct the multi-channel audio signal.

[0005] Specifically, in parametric multichannel audio coding, instead of traditional mid-side or intensity coding, the multichannel input signal is downmixed to a smaller number of channels (e.g., downmixed from two channels to one), and multichannel image (stereo) parameters are extracted. The downmixed signal is then encoded using a more conventional audio encoder (e.g., a single-channel audio encoder). The downmixed bitstream is multiplexed with the encoded multichannel image parameter bitstream. This bitstream is then transmitted to the decoder, where the reverse process is performed. First, the downmixed audio signal is decoded, and then the multichannel audio signal is reconstructed using the encoded multichannel image / upmix parameters.

[0006] Schuijers, W. Oomen, B. den Brinker, and J. Breebarart describe an example of stereo coding in "Advances in Parametric Coding for High-Quality Audio" (114th AES Conference, Amsterdam, Netherlands, 2003, preprint 5852). In the described method, the downmixed single-channel signal is parameterized by utilizing the natural separation of the signal into three components (objects): transient, sine wave, and noise. Further details are provided in "Low Complexity Parametric Stereo Coding" by E. Schuijers, J. Breebarart, H. Purnhagen, and J. Engdegird (116th AES Conference, Berlin, Germany, 2004, preprint 6073), which describes how parametric stereo can be implemented with low (decoder) complexity when combined with frequency band replication (SBR).

[0007] In the described method, decoding is based on the use of a so-called decorrelation process. This decorrelation process generates a decorrelated auxiliary signal from the single-channel signal. During stereo reconstruction, the upmixed stereo signal is generated using both the single-channel signal and the decorrelated auxiliary signal, based on upmixing parameters. Specifically, the two signals can be multiplied by a time- and frequency-dependent factor. The matrix provides the output stereo signal, and the matrix has coefficients determined according to the upmixing parameters.

[0008] However, while parametric stereo (PS) and similar downmixing encoding / decoding methods represent a leap forward compared to traditional stereo and multichannel encoding, they are not optimal in all scenarios. Specifically, known encoding and decoding methods tend to introduce distortions, variations, artifacts, etc., which can introduce differences between the (raw) multichannel audio signal supplied to the encoder and the multichannel audio signal reconstructed at the decoder. Typically, audio quality may degrade and the reconstruction of the multichannel audio signal may not be perfect. Furthermore, the data rate can still be higher than desired and / or the processing complexity / resource usage can be higher than optimal.

[0009] Another issue is that reducing complexity and computational load is often particularly desirable, especially on the decoder side.

[0010] Therefore, improved methods will be advantageous. Specifically, methods that achieve one or more of the following effects will be advantageous: increased flexibility, improved adaptability, enhanced performance, improved audio quality, improved audio quality trade-off with data rate, reduced complexity and / or resource usage, improved input from the encoder side to the decoder side for operation / processing, reduced computational load, convenient implementation methods, and / or improved spatial audio experience. Summary of the Invention

[0011] Therefore, the present invention aims to mitigate, alleviate or eliminate one or more of the above-mentioned disadvantages, preferably alone or in any combination.

[0012] According to one aspect of the present invention, an audio apparatus for generating an output multi-channel audio signal is provided, the audio apparatus comprising a receiver and an audio signal generator. The receiver is arranged to receive an audio data signal comprising (data describing the following): (i) a downmixed audio signal, which is a downmix of a first multi-channel audio signal; (ii) a set of upmixing parameters for time-frequency segments of the downmixed audio signal, each set of upmixing parameters comprising at least: a level difference parameter indicating the level difference between channels of the first multi-channel audio signal (in the time-frequency segment); a correlation parameter indicating the coherence between channels of the first multi-channel audio signal; and a phase difference parameter indicating the phase difference between channels of the first multi-channel audio signal; and (iii) neural network control data comprising: at least one transient parameter indicating the transient characteristics of the first multi-channel audio signal, the transient parameter having a different time-frequency resolution than the set of upmixing parameters. The audio signal generator is configured to generate the output multi-channel audio signal by upmixing the downmixed audio signal according to the upmixing parameter set and the neural network control data. The audio signal generator includes an artificial neural network having input nodes that receive the upmixing parameter set and input nodes that receive transient parameters.

[0013] In many embodiments, this method can provide an improved audio experience. For many signals and scenarios, this method can provide improved generation / reconstruction of multi-channel audio signals with improved perceived audio quality. This method not only provides a significantly improved representation of the transients of the original first multi-channel audio signal in the generated output multi-channel audio signal, but also provides improved upmixing by utilizing parameters that more closely represent the relationships between the channels of the multi-channel audio signal, thus compensating for the effects of transient behavior.

[0014] This method can provide a particularly advantageous arrangement that, in many embodiments and scenarios, allows for the utilization of the convenience and / or improvement of artificial neural networks in audio processing (typically including multi-channel audio decoding / reproduction). This method allows for the advantageous use of artificial neural networks when generating multi-channel audio signals from downmixed audio signals.

[0015] In many embodiments, this method allows for the generation of improved multi-channel audio signals by enabling the encoder to adapt / modify the processing of the apparatus for generating multi-channel audio signals based on transient characteristics. This method can be implemented or facilitated by providing specific methods that are compatible with and utilize the advantages that can be achieved by employing artificial neural networks as part of the audio processing.

[0016] This method offers an efficient implementation and, in many embodiments, allows for reduced complexity and / or resource usage. In many scenarios, this method allows for reducing the data rate representing multi-channel audio signals by using down-mixing. In fact, in many embodiments, significantly improved audio quality can be achieved with a very small increase in the overall data rate.

[0017] In many embodiments, this method can allow for reduced complexity and / or resource usage on the decoder / generation / reconstruction side.

[0018] The samples of the downmixed audio signal can be time-domain samples or frequency-domain samples (specifically, sub-band samples). The samples can span a specific time and frequency range.

[0019] Artificial neural networks are trained artificial neural networks.

[0020] An artificial neural network can be a trained artificial neural network trained on training data, which includes: a training downmixed audio signal generated from a training multichannel audio signal, training upmixing parameters, and neural network control data. Training employs a cost function that compares the training multichannel audio signal with an output multichannel signal generated by an audio signal generator based on the training downmixed signal using the training upmixing parameters and neural network control data. The artificial neural network can also be one or more trained artificial neural networks trained on training data, which includes training data representing a series of related audio sources, including recordings of video, film, telecommunications, etc.

[0021] The generator can be configured to generate a multi-channel audio signal by applying matrix multiplication to the downmixing signal and the auxiliary audio signal, wherein the coefficients of the matrix are determined as a function of the upmixing parameters. The matrix can be time- and frequency-dependent.

[0022] The audio device can specifically be an audio decoding device.

[0023] Transient parameter data values ​​can be referred to as metadata, conditional features, latent representations, conditional variables, and / or composite values.

[0024] Each time-frequency segment may include a time interval and a frequency interval, or correspond to a time interval and a frequency interval. Each time-frequency segment may be a time-frequency block. Time-frequency segments may be disjoint and may be disjoint in both the frequency domain and / or the time domain.

[0025] Time-frequency segments or blocks can be different time intervals and frequency intervals. Each time-frequency segment / block can represent a frequency interval within a time interval. In many embodiments, the first multi-channel audio signal can be divided into time segments / intervals, and the frequency representation of the signal within a time segment / interval can be provided by signal values ​​representing different frequency segments of the signal within the time segment / interval.

[0026] The transient parameters have a different time-frequency resolution than the set of upmixing parameters; this can be a different time resolution and / or a different frequency resolution. In many embodiments, the transient parameters and upmixing parameters can have the same time resolution but different frequency resolutions. In particular, the transient parameters can typically have a coarser frequency resolution than the upmixing parameters. For at least some frequency intervals for which the upmixing parameters provide individual values, the transient parameters can provide only a single parameter value. In many cases, the transient parameters can provide only a single value for the entire frequency band. Therefore, in some embodiments, the spectrum is not divided, and the time-frequency segments / blocks can be time segments / intervals (for neural network control data).

[0027] According to an optional feature of the invention, the audio signal generator is arranged to generate the output multichannel audio signal by applying an upmixing coefficient to the downmixed audio signal and a decorrelation signal generated based on the downmixed audio signal, and the artificial neural network is arranged to generate the upmixing coefficient.

[0028] This can provide advantageous methods for many scenarios and can offer implementations well-suited for employing artificial neural networks, including, for example, providing favorable trade-offs between complexity, computational resources, and / or the perceived audio quality of the generated multichannel audio signal. The audio signal generator can be specifically arranged to generate an output multichannel audio signal by applying matrix multiplication to samples of the downmixed audio signal and the decorrelated signal, where the matrix coefficients are determined by the artificial neural network.

[0029] According to an optional feature of the invention, the audio signal generator is arranged to: generate a decorrelation signal based on the downmixed audio signal, and generate at least one channel of the output multi-channel audio signal by upmixing the downmixed audio signal and the decorrelation signal, wherein the artificial neural network is arranged to control the generation of the decorrelation signal.

[0030] This can provide advantageous methods for many scenarios and can offer implementations well-suited to employing artificial neural networks, including, for example, providing favorable trade-offs between complexity, computational resources, and / or the perceived audio quality of the generated multi-channel audio signals. In many embodiments, the artificial neural network can be arranged to generate signal samples of the decorrelated signal. In some embodiments, the artificial neural network can be arranged to generate parameter values ​​for a decorrelator, to which a downmixed audio signal is applied to generate the decorrelator to produce the decorrelated signal.

[0031] According to an optional feature of the invention, the artificial neural network includes inputs of segments of samples of the downmixed audio signal and outputs of samples of segments of the output multichannel audio signal.

[0032] This can provide advantageous approaches for many scenarios and can offer implementation methods that are well-suited for employing artificial neural networks, including, for example, providing a favorable trade-off between complexity, computational resources, and / or the perceived audio quality of the generated multi-channel audio signals.

[0033] According to an optional feature of the invention, the neural network control data includes the inter-channel level difference for each of a plurality of transients.

[0034] This can provide particularly advantageous neural network control data, which can offer especially suitable information about relevant transient characteristics. In many embodiments, this feature can allow for an improved trade-off between audio quality and data rate.

[0035] According to an optional feature of the invention, the neural network control data includes timing parameters indicating at least one transient timing.

[0036] This can provide particularly advantageous neural network control data, which can offer especially suitable information about relevant transient characteristics. In many embodiments, this feature can allow for an improved trade-off between audio quality and data rate.

[0037] According to an optional feature of the invention, the neural network control data does not include at least some transient inter-channel correlation or inter-channel phase difference data for the first multi-channel audio signal.

[0038] This can provide particularly advantageous neural network control data, which can offer especially suitable information about relevant transient characteristics. In many embodiments, this feature can allow for an improved trade-off between audio quality and data rate.

[0039] According to an optional feature of the invention, the neural network control data has a lower frequency resolution than the upmixing parameters.

[0040] This can provide particularly advantageous operation and often allows for a better trade-off between audio quality and data rate. In many embodiments, at least one transient parameter may not be frequency-dependent. For example, transient parameter values ​​may apply to the entire frequency range of the downmixed audio signal / multichannel audio signal. In some embodiments, transient parameter values ​​may be for multiple sub-bands and may be shared by all sub-bands.

[0041] According to an optional feature of the invention, the neural network control data includes data indicating the transient probability distribution characteristics of the first multi-channel audio signal.

[0042] This can provide advantageous operation and / or implementation methods and / or performance in many embodiments.

[0043] According to one aspect of the present invention, an audio apparatus for generating an audio data signal is provided, the audio apparatus comprising a receiver, a downmixer, a transient detector, and a generator. The receiver receives a first multi-channel audio signal. The downmixer is arranged to: downmix the first multi-channel audio signal into a downmixed audio signal, and determine a set of upmixing parameters for time-frequency segments of the downmixed audio signal, each set of upmixing parameters including at least: a level difference parameter indicating the level difference between channels of the multi-channel audio signal; a correlation parameter indicating the coherence between channels of the multi-channel audio signal; and a phase difference parameter indicating the phase difference between channels of the multi-channel audio signal. The transient detector is arranged to determine at least one transient parameter indicating transient characteristics of the first multi-channel audio signal, the at least one transient parameter having a different time-frequency resolution than the set of upmixing parameters. The generator is arranged to generate the audio data signal including the downmixed audio signal, the set of upmixing parameters, and neural network control data including the at least one transient parameter.

[0044] According to an optional feature of the invention, the transient detector is arranged to: detect a transient in response to detecting that a first level difference metric indicating the level difference between channels of the first multi-channel audio signal differs from a second level difference metric indicating the level difference between the channels by more than a threshold, the first level difference metric being determined for a time interval shorter than the second level difference; and determine the at least one transient parameter to indicate the first level difference metric.

[0045] This can provide advantageous operation and / or implementation methods and / or performance in many embodiments.

[0046] According to one aspect of the present invention, a method for generating an output multichannel audio signal is provided, the method comprising: (I) receiving an audio data signal, the audio data signal comprising: i) a downmixed audio signal, which is a downmix of a first multichannel audio signal; ii) an upmixing parameter set for time-frequency segments of the downmixed audio signal, each upmixing parameter set comprising at least: a level difference parameter indicating the level difference between channels of the first multichannel audio signal; a correlation parameter indicating the coherence between channels of the first multichannel audio signal; and a phase difference parameter indicating the phase difference between channels of the first multichannel audio signal; and iii) neural network control data comprising: at least one transient parameter indicating the transient characteristics of the first multichannel audio signal, the at least one transient parameter having a time-frequency resolution different from the upmixing parameter set; and (II) generating the output multichannel audio signal by upmixing the downmixed audio signal according to the upmixing parameter set and the neural network control data, the output multichannel audio signal being generated based on the output of an artificial neural network having an input node receiving the upmixing parameter set and an input node receiving the transient parameter.

[0047] According to one aspect of the present invention, a method for generating an audio data signal is provided, the method comprising: (i) receiving a first multi-channel audio signal; (ii) downmixing the first multi-channel audio signal into a downmixed audio signal, and determining an upmixing parameter set for time-frequency segments of the downmixed audio signal, each upmixing parameter set comprising at least: a level difference parameter indicating the level difference between channels of the multi-channel audio signal; a correlation parameter indicating the coherence between channels of the multi-channel audio signal; and a phase difference parameter indicating the phase difference between channels of the multi-channel audio signal; (iii) determining at least one transient parameter indicating the transient nature of the first multi-channel audio signal, the at least one transient parameter having a different time-frequency resolution than the upmixing parameter set; and (iv) generating the audio data signal to include the downmixed audio signal, the upmixing parameter set, and neural network control data including the at least one transient parameter.

[0048] These and other aspects, features and advantages of the invention will become apparent and will be illustrated with reference to one or more embodiments described below. Attached Figure Description

[0049] Embodiments of the invention will be described by way of example only with reference to the accompanying drawings, wherein...

[0050] Figure 1 The illustration shows some elements of an example audio device according to some embodiments of the present invention;

[0051] Figure 2 The illustration shows some elements of an example audio device according to some embodiments of the present invention;

[0052] Figure 3 An example of the structure of an artificial neural network is illustrated.

[0053] Figure 4 An example of a node in an artificial neural network is illustrated.

[0054] Figure 5 An example of the transient representation of an audio signal frame is illustrated;

[0055] Figure 6 An example of the transient representation of an audio signal frame is illustrated;

[0056] Figure 7 Examples of elements of a transient detector according to some embodiments of the present invention are illustrated;

[0057] Figure 8 An example of the transient representation of an audio signal frame is illustrated;

[0058] Figure 9 An example of the transient representation of an audio signal frame is illustrated;

[0059] Figure 10 Examples of stereo audio signals, stereo transient audio signals, and stereo residual audio signals are illustrated.

[0060] Figure 11 The illustration shows an example of the arrangement of elements used to train a neural network;

[0061] Figure 12 An example of the transients of an audio signal frame is illustrated;

[0062] Figure 13 An example of the structure of an artificial neural network is illustrated;

[0063] Figure 14 The illustration shows an example of the time-frequency representation (spectral graph) of a stereo applause signal;

[0064] Figure 15 The illustration shows an example of the time-frequency representation (spectral graph) of a stereo applause signal;

[0065] Figure 16 The illustration shows an example of a time segment including a transient stereo signal; and

[0066] Figure 17 The illustration shows some elements of a processor for implementing an audio device according to some embodiments of the present invention. Detailed Implementation

[0067] Figure 1 The illustration shows some elements of an audio apparatus for generating multi-channel audio signals according to some embodiments of the present invention. Figure 2 An example of an audio device is illustrated, which is arranged to generate audio data signals representing multi-channel audio signals (hereinafter referred to as the first multi-channel audio signal). Figure 2 The audio data signal generated by the audio device can be specifically fed to Figure 1 An audio device that can be arranged to generate an output multichannel audio signal as a copy of a first multichannel audio signal. Figure 1 The audio device will also be called a decoder audio device (or simply a decoder), and Figure 2 The audio device will also be referred to as an encoder audio device (or simply an encoder).

[0068] The decoder audio device includes a receiver 101 arranged to receive a data signal / bitstream comprising a downmixed audio signal, which is a downmixed multi-channel audio signal. Specifically, the data signal / bitstream may be generated by the encoder audio device to represent the data signal / bitstream of the first multi-channel audio signal.

[0069] The following description will focus on the case where the multi-channel audio signal is a stereo signal and the downmix signal is a single-channel signal. However, it should be understood that the methods and principles described are equally applicable to multi-channel audio signals with more than two channels and downmix signals with more than one channel (although fewer channels than multi-channel audio signals).

[0070] In addition to the downmixed audio signal, the received data signal includes upmixing parameter data, which comprises a set of upmixing parameters used to upmix the downmixed audio signal. Specifically, upmixing parameters can be parameters indicating the relationship between signals of different audio channels of a multi-channel audio signal (specifically a stereo signal) and / or between the downmixed signal and the audio channels of the multi-channel audio signal. Typically, upmixing parameters can indicate time difference, phase difference, level / intensity difference, and / or measures of similarity (such as correlation).

[0071] The set of upmixing parameters should include at least the following items:

[0072] The level difference parameter indicates the level difference between channels (and specifically two channels) of the first multi-channel audio signal. Specifically, the level difference parameter can be, for example, the interaural intensity difference (IID) and / or interaural level difference (ILD) known from ISO / IEC 23003-3:2020 "Information technology—MPEG audio technology—Part 3: Unified speech and audio coding".

[0073] The correlation parameter indicates the coherence between the channels (specifically two channels) of the first multi-channel audio signal. The correlation parameter can be, for example, the inter-channel cross-correlation (ICC) parameter known from ISO / IEC 23003-3:2020 "Information technology—MPEG audio technology—Part 3: Unified speech and audio coding".

[0074] A phase difference parameter indicates the phase difference between channels (specifically two channels) of a first multi-channel audio signal. Specifically, the phase difference parameter can be, for example, the inter-channel phase difference (IPD), total phase difference (OPD), or channel phase difference (CPD) parameter known from ISO / IEC 23003-3:2020 "Information technology—MPEG audio technology—Part 3: Unified speech and audio coding".

[0075] The parameter set may specifically include the IID, ICC, and IPD parameters as defined in ISO / IEC 23003-3:2020 "Information technology—MPEG audio technology—Part 3: Unified speech and audio coding". In particular, the upmixing parameters IID, ICC, and IPD can be defined as follows:

[0076] Where l and r represent the signal values ​​of the two channels / signals of the first multi-channel audio signal (specifically, the left and right channel signals of the stereo signal), and<a,b> This represents the complex inner product between vectors a and b.

[0077] Typically, upmixing parameters are provided based on each time and each frequency (time-frequency block). For example, new parameters can be provided periodically for each subband in the subband set.

[0078] The encoder audio device is accordingly arranged to receive a first multi-channel audio signal and generate an audio data signal representing the first multi-channel audio signal, the representation including a downmixed audio signal and a set of upmixing parameters. In particular, the encoder audio device may be a parametric stereo (PS) encoder that receives a stereo signal and encodes it into a single-channel audio signal with associated upmixing parameter data.

[0079] Typically, the downmixed audio signal is encoded, and the receiver 101 is arranged to decode the downmixed audio signal to provide the downmixed audio signal (i.e., the single-channel signal in a particular example) along with the set of upmixing parameters and any other desired data.

[0080] Receiver 101 is coupled to audio signal generator 103, which generates a multi-channel audio signal based on a downmixed signal and upmixing parameters. Audio signal generator 103 includes an artificial neural network 105 coupled to a multi-channel audio signal generator 107, which provides the output multi-channel audio signal. Artificial neural network 105 can generate output samples / output values ​​based on upmixing parameters provided to it as input values. In many embodiments, downmixed audio samples may also be provided to artificial neural network 105. Multi-channel audio signal generator 107 is arranged to generate the output multi-channel audio signal based on the output of artificial neural network 105 (and in many cases, also based on the downmixed audio signal). It should be understood that the specific functionality of multi-channel audio signal generator 107 and artificial neural network 105 (and their training) may vary in different embodiments, and various methods will be described later.

[0081] The encoder audio device includes a receiver 201 that receives a first multi-channel audio signal from an internal or external source. The receiver 201 is coupled to a downmixer 203, which is arranged to downmix the first multi-channel audio signal to generate a downmixed audio signal, which is a signal having fewer channels than the first multi-channel audio signal. In addition to the downmixed audio signal, the downmixer 203 continues to generate multiple sets of upmixing parameters, each set of upmixing parameters as previously described with respect to the audio data signal including at least: a level difference parameter indicating the level difference between channels of the multi-channel audio signal; a correlation parameter indicating the coherence between channels of the multi-channel audio signal; and a phase difference parameter indicating the phase difference between channels of the multi-channel audio signal.

[0082] It should be understood that many methods for generating such downmixed audio signals and associated upmixing parameters are known, and any method may be used appropriately without departing from the present invention.

[0083] In many embodiments, the first multi-channel audio signal may specifically be a stereo signal, and the upmixing parameters may be generated from samples of the left and right channel signals of the input stereo signal. In such a case, the downmixed audio signal is a single-channel downmixed audio signal.

[0084] The encoder audio device also includes a data signal generator 205 that generates audio data signals to include data representing downmixed audio signals and upmixing parameters.

[0085] In many embodiments, the encoder audio device and the decoder audio device are arranged to perform sub-band processing. In particular, upmixing parameters can be generated for different (frequency) sub-bands of the first multi-channel audio signal and the downmixed audio signal.

[0086] Specifically, receiver 201 or downmixer 203 may include a filter bank arranged to generate a frequency sub-band representation of the downmixed audio signal. Typically, receiver 201 or downmixer 203 may include a filter bank applied to all channels of the first multi-channel audio signal, such that each channel signal is divided into sub-bands. Downmixing can then be performed on a per-sub-band basis, wherein upmixing parameters are determined for each sub-band and a sub-band downmixed signal is generated. Then, in some cases, the sub-band downmixed audio signal may be directly included in the audio data signal as a sub-band downmixed audio signal, or it may be transformed to the time domain to provide a time-domain signal.

[0087] The filter bank can be a group of quadrature mirror filters (QMFs), or it can be implemented, for example, by a fast Fourier transform (FFT). However, it should be understood that many other filter banks and methods for dividing an audio signal into multiple sub-bands are known and can be used. Specifically, the filter bank can be a complex-valued pseudo-QMF group, thereby producing, for example, 32 or 64 complex-valued sub-bands.

[0088] Furthermore, processing is typically performed within time segments. In most embodiments, the first multi-channel audio signal is divided into time intervals / segments by transforming it to the frequency domain / subband domain by applying, for example, FFT or QMF filtering to samples of each signal. For example, each channel of the multi-channel audio signal can be divided into time segments of, for example, 2048, 1024, or 512 samples. These signals can then be processed to generate samples for, for example, 64, 32, or 16 subbands. Thus, a sample set can be determined for each subband of the downmixed audio signal. Furthermore, an upmixing parameter set can be generated for each time segment / interval and frequency interval / subband.

[0089] It should be noted that the number of time-domain samples is not directly coupled to the number of subbands. Typically, for a so-called critical sampling filter bank with N frequency bands, every N input samples will produce N subband samples (one per subband). An oversampling filter bank will generate more output samples. For example, for every N input samples, this oversampling filter bank will generate... Each output sample consists of k consecutive samples from each frequency band.

[0090] Therefore, a set of upmixing parameters can be generated, where each set is provided for the time interval and a given frequency interval, also known as the time block or segment.

[0091] Each set of upmixing parameters may specifically include IID, ICC, and IPD values ​​as described above, and these parameters are thus provided with a given time resolution and a given frequency resolution. The time interval may vary, but in many embodiments it typically has a fixed duration. In some embodiments, the subband size / frequency resolution may also be fixed / constant for all subbands, but in many embodiments, subbands may have different resolutions / sizes. In many embodiments, the filter bank may be arranged to generate subband signals with subbands of equal bandwidth, and in many other embodiments, the filter bank may be arranged to generate subband signals with subbands of different bandwidths. For example, a higher frequency subband may have a higher bandwidth than a lower frequency subband. Furthermore, subbands may be grouped together to form a subband with higher bandwidth.

[0092] Typically, subbands can have bandwidths ranging from 10 Hz to 10,000 Hz.

[0093] As previously described, the audio signal generator 103 includes an artificial neural network 105 that generates a portion of the output multichannel audio signal from the downmixed audio signal and the set of upmixing parameters. The artificial neural network 105 can be arranged, for example, in various embodiments, to: generate upmixing parameter values / weights for the downmixed audio signal, generate decorrelation auxiliary audio signals corresponding to the downmixed audio signal, directly generate upmixed channel signals for the output multichannel audio signal, etc.

[0094] The artificial neural network used in the described function can be a network of nodes arranged in a hierarchical manner, with each node holding a node value. Figure 3 The illustration shows an example of a portion of an artificial neural network.

[0095] The node value of a given node can be calculated to include contributions from some or typically all nodes in the previous layer of the artificial neural network. Specifically, the node value of a node can be calculated as a weighted sum of the node values ​​of all nodes in the previous layer. Typically, biases can be added, and the result can be processed by an activation function. Activation functions typically provide the basic components of each neuron by introducing nonlinearity. Such nonlinearity and activation functions provide significant effects in the learning and adaptation processes of neural networks. Therefore, node values ​​are generated as a function of the node values ​​in the previous layer.

[0096] The artificial neural network may specifically include an input layer 301, which includes multiple nodes that receive input data values ​​from the artificial neural network. Therefore, the node values ​​of the nodes in the input layer can typically be directly the input data values ​​of the artificial neural network, and thus can be calculated without relying on the values ​​of other nodes.

[0097] Artificial neural networks may also include zero, one, or more hidden layers 303, 305, or processing layers. For each of such layers, node values ​​are typically generated as a function of the node values ​​of the previous layer, and in particular, these values ​​are weighted combinations with biases, followed by the application of an activation function (e.g., sigmoid, ReLU, or tanh functions may be applied).

[0098] In particular, such as Figure 3 As shown, each node (also called a neuron) can receive input values ​​(from nodes in the previous layer) and thereby compute node values ​​as a function of these values. Typically, this involves first generating values ​​that are linear combinations of the input values, where each of these values ​​is weighted by a weight:

[0099] Where w refers to the weight, x refers to the node in the previous layer, and n refers to the index of the different node in the previous layer.

[0100] The activation function can then be applied to the resulting combination. For example, the node value l can be determined as:

[0101] This function could be, for example, a Rectified Linear Unit function, as described by Xavier Glorot, Antoine Bordes, and Yoshua Bengio in "Proceedings of the Fourth International Conference on Artificial Intelligence and Statistics" (PMLR 15, pp. 315-323, 2011):

[0102] Other commonly used functions include the sigmoid function or the tanh function. In many implementations, multiple functions can be used to compute node outputs or node values. For example, the ReLU and sigmoid functions can be combined using activation functions such as the following:

[0103] Such operations can be performed by each node of an artificial neural network (except for the input nodes, which are usually excluded).

[0104] The artificial neural network also includes an output layer 307, which provides the output from the artificial neural network; that is, the output data of the artificial neural network is the node values ​​of the output layer. As for the hidden / processing layers, the output node values ​​are generated as a function of the node values ​​of the previous layer. However, unlike the hidden / processing layers, where node values ​​are typically inaccessible or unusable, the node values ​​of the output layer are accessible and provide the results of the artificial neural network's operations.

[0105] Many different network architectures and toolkits have been developed for artificial neural networks, and in many embodiments, artificial neural networks can be based on networks that are adapted and customized. An example of a network architecture that can be applied to the above applications is WaveNet by van den Oord et al., which is described in “Wavenet: A generative model for raw audio” (arXivpreprint arXiv:1609.03499, 2016) by Oord, Aaron van den, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu.

[0106] WaveNet is an architecture for synthesizing temporal signals using dilated causal convolutions and has been successfully applied to audio signals. For WaveNet, the following activation function is typically used:

[0107] Where * denotes the convolution operator, ⊙ denotes the element-wise multiplication operator, σ(∙) is the sigmoid function, k is the layer index, f and g represent the filter and gate, respectively, and W represents the weights of the learned artificial neural network. The filter product of the equation typically provides a filtering effect, while the gate product provides weighting of the result, which in many cases can effectively allow the contribution of a node to be reduced to essentially zero (i.e., it can allow or "cut off" nodes that contribute to other nodes, thus providing a "gate" function). In different cases, the gate function can result in the output of that node being negligible, while in other cases it will make a significant contribution to the output. Such functions can greatly help allow neural networks to learn and be trained efficiently.

[0108] In some cases, artificial neural networks can also be configured to include additional contributions that allow the network to be dynamically adapted or customized for specific desired properties or characteristics of the generated output. For example, a set of values ​​can be provided to adapt the artificial neural network. These values ​​can be included by contributing to some nodes of the network. Specifically, these nodes can be input nodes, but are typically nodes in hidden or processing layers. Such adaptation values ​​can, for example, be weighted and added as contributions to the weighted sum / correlation value of a given node. For example, in WaveNet, such adaptation values ​​can be included in the activation function. For example, the output of the activation function can be given as:

[0109] Where y is a vector representing the fit values, and V represents the appropriate weights for these values.

[0110] The above description relates to neural network methods that can be applied to many embodiments and implementations. However, it should be understood that many other types and structures of neural networks can be used. In fact, many different methods for generating neural networks have been and are being developed, including neural networks using complex structures and processes different from those described above. This method is not limited to any particular neural network method, and any suitable method can be used without departing from the present invention.

[0111] In this method, an encoder audio device is arranged to generate control data for an artificial neural network 105. This control data is included in the audio data signal and transmitted to a decoder audio device, where it is provided as input to the artificial neural network. Therefore, the neural network control data is the data processed by the artificial neural network 105, and the output of the artificial neural network 105 depends on the neural network control data.

[0112] An encoder audio device is arranged to determine neural network control data to include at least one transient parameter that depends on / represents the transient characteristics of a first multi-channel audio signal.

[0113] In different embodiments, the specific transient parameters and characteristics transmitted and input to the artificial neural network 105 may differ. For example, in many cases, transient data may include indications of one or more of the following transients present in the first multi-channel audio signal (and typically in the downmixed audio signal): presence, quantity, time, duration, amplitude, inter-channel level difference, inter-channel phase difference, and inter-channel coherence.

[0114] Transient data / parameters are provided with a different time-frequency resolution than the set of overmixed parameters, and in many cases can be provided with a finer timing resolution or a coarser frequency resolution than the set of overmixed parameters, and in many cases both a finer time resolution and a coarser frequency resolution. For example, the timing of transients can be indicated with a finer time granularity than the (processing) time segments / intervals, and, for example, non-frequency-dependent inter-channel level differences can be provided.

[0115] A set of upmixing parameters can be provided for time-frequency blocks corresponding to subbands and time segments in a process as described above (e.g., each time segment and subband has a fixed number of samples). However, in some cases, transient parameter values ​​can be provided at a finer temporal resolution. For example, the timing or duration of a transient can be provided at a higher temporal resolution than the sampling time of the subband samples.

[0116] Furthermore, in many embodiments, transient characteristics can be provided with a coarser frequency resolution than the set of upmixing parameters. In particular, in some embodiments, the set of upmixing parameters can be provided for each subband, while a single common transient parameter value can be provided for multiple subbands, and possibly all subbands.

[0117] In many embodiments, one or more sets of parameters may be provided for each transient in the set of detected transients. In particular, the parameters may include an indication of the channel level difference between channels of the first multi-channel audio signal.

[0118] Figure 6 A specific example of a time segment / frame FRM from which three transients are detected is shown. Here, for a given time segment / frame, three transients are detected at positions p0, p1, and p2. These positions / moments can be encoded and included in the audio data signal, and thus transmitted to the decoder audio device. Furthermore, for each transient parameter position p, the amplitude level difference of the transient is determined and included in the audio data signal. For example, the IID (interaural intensity difference) can be determined and included in the audio data signal. In such a case, a positive IID may correspond to a left shift of the transient signal, while a negative IID may correspond to a right shift of the transient signal. In some embodiments, the individual amplitudes a0, a1, and a2 may be transmitted additionally or alternatively. Thus, the transient data can be used to encode a representation of the detected transients.

[0119] In some embodiments, such as Figure 6 As shown, the duration of a transient can be determined, encoded, and transmitted to the decoder audio device in the audio data signal.

[0120] In many embodiments, parameter values ​​can be quantized into a relatively small number of levels, and therefore each value can use a relatively small number of bits. In many embodiments, the word length can be no more than 1, 2, 3, or 4 bits. For example, amplitude values ​​can be quantized into several levels (e.g., 5 or 7 discrete levels) using only a few bits.

[0121] The encoding of parameter values ​​in audio data signals can, for example, use absolute or differential encoding. In particular, the three IID values ​​corresponding to positions p0, p1, and p2 can be differentially encoded for the (band-averaged) IID transmitted across the entire frame pair (each band).

[0122] The encoder audio device can correspondingly generate transient data indicating transients in the first multi-channel audio signal and provide this transient data to the decoder audio device, where it is input into the artificial neural network 105. The encoder audio device can provide this transient data at a different time-frequency resolution than the upmixing parameters, thereby allowing independent optimization of the transient data. In particular, a coarser frequency resolution can be used, and in many scenarios, the transient data may not include any frequency dependence, but can provide the same parameter values ​​for all sub-bands of the downmixed audio signal. In many embodiments, the parameter values ​​can also be coarsely quantized to several discrete levels. Therefore, very low data overhead can be achieved in many embodiments. However, it has been found that providing this transient data / information to the artificial neural network 105 of the upmixer can result in significantly improved perceived audio quality, and in particular, can significantly improve the perceived audio realism of some scenes and environments.

[0123] It should be understood that different methods can be used to detect transients in the first multi-channel audio signal and / or the downmixed audio signal (in fact, these operations can be considered equivalent, since transients in the first multi-channel audio signal also exist in the downmix, and therefore detecting transients in the first multi-channel audio signal also detects transients in the downmixed audio signal, and vice versa).

[0124] Figure 7 An example of the elements of a transient detector 207 that can be used in an encoder audio device is illustrated. In this example, the transient detector 207 can be arranged to detect transients in a stereo signal.

[0125] Transient detection can be performed independently for different channels, and specifically for the left and right channels in this example, the detection results are then combined. In other embodiments, information from both channels can be used directly, for example, by considering the downmixed audio signal. Such an approach may be advantageous because it can utilize inter-channel level (e.g., IID) parameters. The following examples will primarily consider this type of approach.

[0126] In many embodiments, as described below, the transient detector 207 can detect a transient in response to the following detection: a first level difference metric (specifically, a first IID metric) indicating the level difference between channels of the first multi-channel audio signal differs from a second level difference metric (specifically, a second IID metric for the same channel) indicating the level difference between channels by exceeding a threshold, wherein the first level difference metric is determined for a time interval shorter than the second level difference. In many embodiments, the shorter time interval may not exceed 10%, 20%, 30%, or 50% of the time interval of the second level difference.

[0127] exist Figure 7 In the example, transient detector 207 includes an analysis filter bank 701 that typically decomposes the left and right stereo channels into time-frequency (TF) representations, where the distribution of center frequencies follows, for example, the critical bands of logarithmic intervals in the human auditory system (inner ear). Spectral decomposition can be performed using a hybrid quadrature mirror filter bank (QMF), which produces fine resolution at low frequencies, where the resolution decreases (bandwidth increases) with increasing frequency. It should be understood that in many embodiments, such an analysis filter bank 701 can be equivalently part of receiver 201, and the generated sub-band representation can also be used for downmixing, upmixing parameter estimation, etc.

[0128] QMF decomposition can produce complex-valued outputs, and the transient detector 207 includes envelope circuitry 703 that determines the real envelopes of both the left and right channels. An example of the envelope is given by the following equation:

[0129] in and Is for The time-frequency samples of time slice m and frequency bin k are the real and imaginary parts. Square root operations can be omitted to reduce complexity. This envelope is an (absolute) spectrogram where no summation is performed across different frequency bands to generate the time-domain envelope.

[0130] It should be noted that, according to the embodiments, it is not always necessary to calculate the real envelope, because the IID calculation already includes the calculation of the square of the complex signal amplitude.

[0131] The transient detector 207 includes a detection circuit 705 coupled to an envelope circuit 703 and, in this example, processes temporal frequency samples of the current and previous frames. In this example, the IID on the windowed frame is determined (e.g., using a method similar to that of a conventional PS encoder containing a symmetrical Hanning window). This IID value is used as a baseline to predict the perceptual effect of the detected transient in later processing and is denoted as... .

[0132] A positive baseline value indicates a left shift on the frame, a value close to zero indicates a center shift (no shift), and a negative baseline IID indicates a right shift.

[0133] Use a short sliding window corresponding to approximately 10ms To calculate the IID between the left and right channels:

[0134] Deviation of Values ​​can be considered transient candidates, and these transient candidates can have their values ​​compared to the baseline IID. Separately coded IID values. It should be noted that IID values ​​can be calculated at each QMF time in each frequency band, or aggregated globally or across frequencies according to a custom grouping scheme that may or may not omit certain frequencies irrelevant to the detection of certain transients.

[0135] Next, the perception-driven step can evaluate the perception performance of the (PS) encoder on the detected transients. It can... Sets and It also filters out transients that are not affected by baseline (traditional) IID reconstruction in the decoder. This helps reduce the number of parameters that must be sent in the bitstream, thus keeping the resulting bit rate under control.

[0136] It is known that humans perceive slowly changing stereo parameters as moving sound sources, but can only detect rapidly changing parameters as an increase or decrease in the width of the stereo image.

[0137] The perceptual filtering step has two objectives: 1. Decide whether the IID parameters should be calculated and transmitted separately for a given transient. 2. Group transients together based on their IID characteristics, that is, group transients originating from the same source / location.

[0138] The primary objective is perception-driven and based on the bias between the transient IID parameters and the overall estimated frame parameters. If this bias exceeds a certain threshold, the transient is included as a portion of the stereo transient parameters.

[0139] The second objective is to further reduce the bit rate, but this requires bundling transients based on their stereo characteristics and assigning them to virtual sources (objects) in the stereo image. This way, if the IID value of a given source is stable over time, only timing information needs to be transmitted for the same source, without simultaneously transmitting timing and IID features.

[0140] Figure 8 The diagram illustrates the relationship between... Figure 5 and Figure 6 The same stereo transient representation is shown, but also ( The average IID of the frame and the average IID of the frame. The range of IIDs around the given average IID is given. In such a case, transients falling outside the indicated range can be represented by parameters included in the audio data signal.

[0141] The value can be adjusted based on a perception (listening) test or model, which can, for example, determine the minimum perceptible difference (JND) between the transient frame IID and the average frame IID.

[0142] However, since JND is also known to be a function at the IID level—generally, the larger the IID, the larger the JND— The area between can be used Replace the exclusion area surrounding the level itself, such as Figure 9 As shown. If the average IID of the frame falls within this range, the transient can be ignored and will not be encoded (as in the example transient). (The situation).

[0143] The perceptual filtering step may also include a masking model. Similar to exclusion regions, for example, a masking model can indicate to the user that a given transient cannot be fully perceived based on background noise.

[0144] In the example, short-term characteristics are compared with long-term characteristics accordingly, and transients are adequately detected based on these differences.

[0145] Many suitable transient detection methods are based on tracking the changes in the signal envelope (wideband or per spectrum) relative to slowly changing or the rapid changes in the residual to detect the start (and end) of a transient.

[0146] In another method, the magnitude of the time-frequency representation of the signal (e.g., a channel signal of a first multi-channel audio signal or a downmixed audio signal) is determined, and the resulting frequency envelope is summed across the frequencies. Two smoothed versions of this envelope can then be created, one tracking the envelope more slowly than the other. Thus, the slowly changing residual envelope value is determined, and the other value is determined using the much faster time constant of the tracking envelope. A first-order exponential smoothing or, for example, a smoothed moving average filter can be applied to create the smoothed envelope. In the case of a fast-tracking envelope, an instantaneous envelope can also be used.

[0147] in and These are slow and fast envelopes, respectively. and .

[0148] Then, the ratio between the fast tracking envelope and the slow tracking envelope can be used to indicate the abrupt change in the transient relative to the residual signal, and can be used as a time-domain (wideband) gain function.

[0149] As an indication of the existence of transients, ,and Such methods are described, for example, in (Adami, A., Herzog, A., Disch, S., and Herre, J., 2017, October) "Transient-to-noise ratiorestoration of coded applause-like signals" (2017 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), pp. 349-353, IEEE).

[0150] This metric can be used to detect transients and their timing. Specifically, if g(n) exceeds a threshold, a transient can be considered to have been detected. In some cases, the metric g(n) can then be compared with a second threshold (specifically, it can be the same as the first threshold), and if the metric falls below the second threshold, the end of the transient can be considered to have been detected. Thus, this method can detect the start and end of a transient, and correspondingly, the duration.

[0151] The interval between the start and end of a transient in the first multi-channel audio signal and / or downmixed audio signal can also be used to filter out non-transient components with longer durations, and thus transient detection can detect only relatively short transients rather than step changes with longer durations.

[0152] In some embodiments, the encoder audio device may, in certain cases, separate the downmixed audio signal into a set of transients and remove the residual signal of these transients. For example, a portion of the downmixed audio signal between the detection of the start and end of a transient may be extracted and represented as a separate transient, while the resulting downmixed audio signal represents the residual signal.

[0153] In some cases, a softer separation of the signal into transient and residual signals can be performed by using weighted selection instead of binary selection. For example, the transient signal t(n) can be generated by multiplying the first multi-channel audio signal and / or the down-mixed audio signal by the detection signal g(n), i.e.

[0154] Where m(n) represents, for example, the channel signal of the first multi-channel audio signal or the downmixed audio signal.

[0155] Similarly, residual signals can be generated, for example, as follows:

[0156] Figure 10 The diagram illustrates the separation of stereo transient signals. and stereo residual signal stereo input signal (An example where two channels are represented as overlapping each other and in different grayscale values.)

[0157] Another method for separating transients is to track the residual signal by frequency bandwidth using minimum tracking of the envelope (based on, for example, the minimum statistics method for tracking stationary noise described in Martin, Rainer, “Noise power spectral density estimation based on optimal smoothing and minimum statistics” (IEEE Transactions on speech and audio processing 9.5, 2001, pp. 504-512). .

[0158] In this case, the transient signal can be written in the frequency domain as:

[0159] in It is a frequency domain estimate of the residual signal using minimum tracking, and It is an over-subtraction factor (≥1.0) used to compensate for the underestimation of the residual signal.

[0160] As another example, source separation techniques using neural networks can be employed. Examples of such techniques can be found in "Wave-U-Net: A Multi-Scale Neural Network for End-to-End Audio Source Separation" by Daniel Stoler, Sebastian Ewert, and Simon Dixon (http: / / arxiv.org / abs / 1806.03185, 2018).

[0161] Transient detection can be performed using a neural network model that has been trained to detect the location of transients using various neural network implementations, such as fully connected layers, convolutional layers, or recurrent layers. Figure 11 The diagram below illustrates a block diagram of a neural network and its training.

[0162] In this example, the input corresponds to the current frame or a combination of the current frame (F) and the previous frame (F-1), where the time block of the previous frame can be used as additional padding if the transient occurs at the beginning of the current block F. Furthermore, the samples can correspond to all time-frequency samples or a range of frequency samples relevant to transient detection (e.g., for clapping, most of the energy is between 1 and 3 kHz). Assuming the use of... If stereo data blocks are used as input, then the training data can be derived from... The input blocks consist of, and with The 1 or 0 vectors correspond to labels indicating transient locations in the training data. Additionally, transient presence flags indicating whether a frame includes at least a single transient can be added to the training labels, thereby generating... Vector labels.

[0163] The loss function used can correspond to aggregate cross-entropy loss (assuming a frame length of n samples):

[0164] in Corresponding to a given frame i th The baseline truth label for the time instance, and It is the neural network at time 1000. i There are transient prediction probabilities.

[0165] Another loss function can be based on the mean squared error between the baseline true value and the predicted label.

[0166] Manually labeling foreground clapping sounds for applause signals is possible, although time-consuming. Therefore, for the purposes of this invention, training can be performed on synthetic data using several available transient synthesizer models (e.g., clapping / applause). These models can generate realistic-sounding clapping signals using signal processing techniques. During inference, the model estimates the location of the detected stereo transient, which can then be used to calculate the corresponding IID value.

[0167] As previously described, the audio signal generator 103 can generate output multi-channel audio signals by using an artificial neural network as part of the processing, wherein the artificial neural network has inputs including both a set of upmixing parameters and transient parameters. Different methods can be used in different embodiments, and in particular, in different embodiments, different signals or parameters can be generated by the artificial neural network for use as part of the upmixing / multi-channel audio signal generation.

[0168] In some embodiments, the artificial neural network 105 includes input to a sample segment of the downmixed audio signal. In some embodiments, the sample may be a time-domain sample of the downmixed audio signal. In other embodiments, the sample may be, for example, a time-frequency sample, such as, in particular, a sub-band sample of a frame / segment.

[0169] In some cases, the artificial neural network 105 is arranged to directly generate output multi-channel audio signals. Therefore, the artificial neural network 105 may have output nodes that directly provide samples of the output multi-channel audio signals. In some embodiments, the samples may be directly time-domain samples of the channel signals of the output multi-channel audio signals. In other examples, the artificial neural network 105 may generate sub-band samples of the output multi-channel audio signals. The generated output multi-channel audio signal samples may be fed to a multi-channel audio signal generator 107, which, in the previous example, may simply forward the samples or, for example, perform post-processing on the generated samples of the output multi-channel audio signals. In the latter case, the multi-channel audio signal generator 107 may perform processing, for example, including frequency-domain to time-domain transformation, to convert the sub-band samples into time-domain samples.

[0170] In some such cases, the decoder audio device can be arranged to generate a decorrelated version of the downmixed audio signal by inputting the downmixed audio signal to a suitable decorrelator. In such cases, samples of both the downmixed audio signal and the decorrelated signal can be, for example, provided to an artificial neural network 105 to generate samples of the output multichannel audio signal.

[0171] In such methods, training of the artificial neural network 105 can be performed, for example, by generating a large number of training input multichannel audio signals, which are processed by an encoder audio device according to, for example, a specified specification or standard. The data can then be provided to a decoder audio device, which can generate output multichannel audio signals based on these data. The cost function used for training can then be determined by comparing the generated output multichannel audio signals with the original training input multichannel audio signals.

[0172] In some embodiments, the decoder audio device may be arranged to generate an output multi-channel audio signal in response to an upmixing of a downmixed audio signal and a decorrelation signal (as a decorrelation version of the downmixed audio signal). The decorrelation signal may have the same overall characteristics as the downmixed audio signal in terms of frequency envelope, average energy, etc., but is decorrelationed with the downmixed audio signal.

[0173] For example, for a traditional parametric stereo PS, upmixing of a single-channel downmixed audio signal m can be based on the decorrelation signal d by using the following upmixing method:

[0174] This upmixing process typically operates in a time- and frequency-dependent manner, where the upmixing parameters are time- and frequency-dependent. For example, a set of upmixing coefficients is usually determined for each time segment / frame and each subband. h xy The upmixing coefficients are determined based on the received set of upmixing parameters. The exact dependence of the upmixing coefficients on the set of upmixing parameters will depend on the specific implementation. For example, in many implementations, the relationships defined for the traditional PS in ISO / IEC 23003-3:2020 "Information technology—MPEG audio technology—Part 3: Unified speech and audio coding" can be used:

[0175] In some embodiments, the artificial neural network 105 can be arranged to generate decorrelated signals, and thus, based on inputs including samples of downmixed audio signals, a set of upmixing parameters, and one or more transient parameters, the artificial neural network 105 continues to generate samples of decorrelated signals. These samples (and thus the decorrelated signals generated by the artificial neural network 105) can then be fed to an audio signal generator 103, in which matrix multiplication upmixing is performed.

[0176] In some embodiments, the audio signal generator 103 is arranged to generate an output multi-channel audio signal from the downmixed audio signal and from the decorrelated audio signal based on parameterized upmixing data. For stereo cases, the generator may specifically use time- and frequency-dependent... Matrix multiplication is applied to samples of the downmixed audio signal and the decorrelated signal to generate the output multichannel audio signal. It is typically based on time and frequency band, and determined from the upmixing parameter data. The coefficients of the matrix. For other upmixing operations, such as from single-channel or stereo downmixing signals to five-channel multi-channel audio signals, the audio signal generator 103 can apply matrix multiplication of matrices with appropriate dimensions.

[0177] It should be understood that those skilled in the art will recognize many different methods for generating such a multi-channel audio signal from the downmixed audio signal and the decorrelation signal, and for determining appropriate matrix coefficients based on the upmixing parameter data, and any suitable method can be used. In particular, those skilled in the art are familiar with various methods for PS upmixing based on the downmixed and auxiliary audio signals.

[0178] In traditional systems, upmixing involves generating a decorrelated signal of a single-channel audio signal, determined by applying the downmixed audio signal to a decorrelation function. It has been found that by generating the decorrelated signal and mixing it with the single-channel audio signal, an improved quality upmix signal can be perceived, and decoders have been developed to take advantage of this. The decorrelated signal is typically generated by a decorrelation function in the form of an all-pass filter applied to the single-channel audio signal. However, while the use of such an all-pass filter tends to result in a multi-channel audio signal perceived as having improved quality, it is still not ideal, and some audio quality degradation can often be perceived.

[0179] In some embodiments, the decorrelation signal is not generated by direct filtering of the downmix / single-channel audio signal, but is generated by an artificial neural network 105, wherein the multi-channel audio signal generator 107 uses the decorrelation signal to generate a multi-channel audio signal based on the upmix parameters.

[0180] Therefore, in some embodiments, the artificial neural network 105 can be directly trained to generate decorrelated signals. It has been found that in many embodiments, this can provide significantly improved performance, typically resulting in output multi-channel audio signals that sound more realistic.

[0181] In some embodiments, the artificial neural network 105 may not directly generate the decorrelated signal, but may instead generate, for example, one or more parameters for a decorrelist that is applied to the downmixed audio signal to generate the decorrelated signal. For example, the artificial neural network 105 may generate parameters (e.g., filter coefficients) as output for an all-pass filter applied to the downmixed audio signal to generate the decorrelated signal.

[0182] Similar to the previously described method, the training of the artificial neural network 105 can be performed, for example, by generating a large number of training input multi-channel audio signals, which are processed by an encoder audio device according to, for example, a specified specification or standard. The generated data can then be provided to a decoder audio device, which can generate output multi-channel audio signals based on these data. The cost function used for training can then be determined by comparing the generated output multi-channel audio signals with the original training input multi-channel audio signals. Therefore, the training of the artificial neural network 105 can be based on an end-to-end cost function, which includes upmixing, etc. It has been found that such trained artificial neural networks produce improved audio quality in many scenarios and for many signals.

[0183] In some embodiments, the artificial neural network 105 can be arranged to directly generate upmixing coefficients (such as those specifically for upmixing the downmixed audio signal and one or more auxiliary signals) for upmixing the downmixed audio signal and one decorrelation signal. In the example above, the artificial neural network 105 can therefore directly generate the coefficients of the upmixing matrix:

[0184] In this method, the upmixing process may include a pre-trained network that determines upmixing parameters applied to the downmixed audio signal and at least one decorrelation signal to generate an upmixed output multichannel audio signal.

[0185] In this method, the artificial neural network 105 can be trained to directly generate coefficients h for a given set of conventional PS parameters (represented by the set of overmixing parameters) and a given set of stereo transient sequences (represented by transient parameters). xx This minimizes the (perceptual) loss of the synthesized stereo output signal compared to the original stereo signal. An example architecture could be a fully connected network that receives the PS parameters of the current frame, the PS parameters of the previous frame, the current sample index n, and the value of the stereo transient sequence at that location. Figure 12 The diagram illustrates an example of the time relationship between input parameters.

[0186] Figure 13 The diagram illustrates simplified graph elements of a fully connected network with input nodes for the PS parameters, the stereo transient sequence s[n], and the relative sampling position n within a frame. For simplicity, only a few connections and the real-valued part of the mixing term h are shown.

[0187] In some embodiments, instead of feeding the artificial neural network 105 (conventional) PS parameters of the previous frame F-1 and the current frame F, the hypermixing terms at frames F-1 and F can be pre-computed using conventional hypermixing equations. In some cases, such coefficients, calculated by predetermined formulas / equations, can then be input into the artificial neural network 105, and modified coefficients can be computed. This can result in a less complex network because network capacity does not need to be applied when modeling the original hypermixing equations.

[0188] Alternative network architectures can employ so-called gated activation units, such as those used in WaveNet. A gated activation unit can be described as:

[0189] Where x is the input signal (at a certain layer / location in the network), and s is the stereo transient signal. W f It is a filter (convolution) applied to the input signal (at a certain layer / location in the network). W g It is a gate (convolution) applied (at a certain layer / location in the network) to the stereo transient signal. It is a nonlinear sigmoid function, and tanh() is a nonlinear tangent function. When the stereo transient is zero, the combination of all the left-hand parts in the above equation will effectively simulate traditional stereo upmixing, while when the gate function is activated, the combination of the left-hand and right-hand parts in the above equation will take into account the stereo image parameters to generate an appropriate stereo output.

[0190] In some embodiments, the neural network control data includes data indicating the probability distribution characteristics of the transients of the first multi-channel audio signal. Therefore, the transient data may (alternatively or additionally) provide an indication of the probability distribution of the transients, rather than (only) providing a fully deterministic stereo transient representation. The transient data may include random components, where some parameters, such as amplitude and / or duration, can be synthesized using a random (noise-like) process.

[0191] More specifically, for example, if we assume that the transient amplitude IID parameter follows a Gaussian distribution:

[0192] On the decoder side, we can use a zero-bias, unit-variance Gaussian noise generator. To generate values ​​that follow the same distribution:

[0193] The transient amplitude IID parameters can then be generated as follows:

[0194] This requires determining parameters on the encoder side. and This is then transmitted to the decoder. This can be achieved by simply measuring the mean and variance of the collected transient amplitude IID parameters and then quantizing them into a finite number of discrete values:

[0195] The example of a Gaussian distribution is provided above. Other probability density functions can be applied; for example, a uniform distribution can be used for the time location of the generated transient.

[0196] In the described method, an artificial neural network is used for upmixing to generate the output multichannel audio signal. In addition to the normal signals and parameters used to perform upmixing (e.g., downmixing audio signals and upmixing parameters related to the characteristics between different channels of the original multichannel audio signal), the artificial neural network is also provided with control data determined based on the transient characteristics of the original multichannel signal / downmixing audio signal.

[0197] It has been found that, in many scenarios and for many signals (and embodiments), the generation of particularly advantageous upmixed multichannel signals can be achieved. It has been found that including transient data and information allows artificial neural networks to consider information that might otherwise be lost (due to insufficient representation by the upmixing parameters), thus allowing for better training and output of the artificial neural network. Furthermore, it has been found that, by using a representation of transient information with a different time-frequency resolution than that used for the upmixing parameters, significant improvements in upmixing and audio quality can be allowed while introducing only a small amount of overhead. In particular, it has been found that very coarse frequency resolutions (including transient parameter values ​​that are not frequency-dependent) can still allow for very accurate and significantly improved upmixing that includes transient components. Moreover, it has been found that different time resolutions (in particular, including those allowing for finer time resolutions than those used for the upmixing parameters) allow for improved audio quality and, in particular, allow for better representation of some audio components and sounds.

[0198] To illustrate this, consider an audio signal representing the applause of a group of people. Such an applause signal tends to consist of a superposition of seemingly random (spatial) individual clapping sounds (i.e., short bursts of energy in time). On the one hand, estimating a set of stereo parameters and interpolating these parameters between frames does not lead to an accurate reconstruction of the stereo signal at the decoder. On the other hand, fine-grained estimation of transient-specific parameters may also increase the total bit rate.

[0199] To further understand the impact of low (per frame) parameter update rates, consider Figure 14 The image illustrates an example of the time-frequency decomposition (spectral graph) of a clapping stereo signal sampled at 44.1 kHz. The signal consists of background and foreground clapping, with the background clapping dominating below 3 kHz. The foreground clapping is noticeably shifted to the left, as it is not prominently present in the right channel. The time-frequency energy of the background clapping is fairly random and resembles noise.

[0200] Figure 15 The image shows the output of a traditional parametric stereo decoder for the same segment. It should be noted that this approach causes both background applause and foreground clapping to become blurred. Traditional stereo parametric interpolation methods are ineffective here. For the background applause, the blurring effect results in a musical tone-like effect and clearly generates harmonic components. The foreground clapping is also slightly blurred temporally and loses its translational characteristics (IID): some foreground clapping now appears more prominently in the right channel. This latter effect is a consequence of the stereo parameter estimator's frame-by-frame update rate.

[0201] To better understand the impact on stereo parameters, consider Figure 16 The image depicts the left channel spectrogram of two frames of data. The foreground clapping sound appears in the middle of the previous frame. Stereo parameters are estimated by first windowing the two frames, which attenuates the energy near the beginning of the previous frame and near the end of the current frame.

[0202] If IID, IPD, and ICC are calculated on these two frames, the distribution of phase, intensity, and coherence of the background applause from the left and right channels will largely determine the estimated stereo parameters, even though the parameters should be different for the foreground clapping sound, since the signal is clearly shifted to the left, at least for IID.

[0203] Therefore, providing transient data with higher temporal resolution allows additional information to be included in the artificial neural network 105 during the training process, thereby allowing the generation of improved output audio signals.

[0204] This method allows for the adaptation of transient information to specific importance. Indeed, for transients, artificial neural networks have been found to be less sensitive to frequency dependence than to temporal accuracy, and in the described method, transient data can be adapted / generated accordingly, without being limited to following the same resolution as the overmixed data.

[0205] Furthermore, it has been found that in many embodiments, the inter-channel level difference can be significant, in addition to timing information indicating the timing of transients, and allows for, for example, an accurate representation of the spatial location of transients (e.g., in a stereo image).

[0206] In many embodiments, the only inter-channel information provided for transients may be an inter-channel level difference indication. Specifically, in many embodiments, transient data may not include any inter-channel phase difference or inter-channel coherence. In practice, such information provides relatively irrelevant information to the artificial neural network and typically has a significantly small impact on the resulting audio quality of the generated output multichannel signal.

[0207] The artificial neural network 105 is specifically trained to provide suitable output data by employing a training process (e.g., as part of the manufacturing or design phase). The results of the training process (e.g., coefficients of different nodes, etc.) can be performed once and then used for all manufacturing devices.

[0208] Artificial neural networks (ANNs) are adapted for a specific purpose through a training process that adapts / tunes / modifies the weights and other parameters (e.g., biases) of the ANN. It should be understood that many different training processes and algorithms are known for training ANNs. Typically, training is based on a large training set, where a large number of examples of input data are fed to the network. Furthermore, the output of the ANN is usually (directly or indirectly) compared to an expected or desired result. A cost function can be generated to reflect the desired outcome of the training process. In a typical scenario known as supervised learning, the cost function typically represents the distance between the prediction and the ground truth value of a particular input data set. Based on the cost function, the weights can be changed, and by repeating the process with the modified weights, the ANN can be adapted to a state that minimizes the cost function.

[0209] More specifically, during the training phase, a neural network can have two distinct information flows: from input to output (forward propagation) and from output to input (backward propagation). In forward propagation, the data is processed by the neural network as described above, while in backpropagation, the weights are updated to minimize the cost function. Typically, such backpropagation follows the gradient direction of the cost function surface. In other words, by comparing the predicted output with a baseline of true values ​​from a set of data inputs, the direction of minimizing the cost function can be estimated and backpropagated by updating the weights accordingly. Other known methods for training artificial neural networks include, for example, the Levenberg-Marquardt algorithm, the conjugate gradient method, and Newton's method.

[0210] In the current context, training can specifically include a training set comprising a potentially large number of multi-channel audio signals. In some embodiments, the training data can be multi-channel audio signals in time segments corresponding to processing time intervals of the artificial neural network being trained; for example, the number of samples in the training multi-channel audio signals can correspond to the number of samples corresponding to the input nodes of the artificial neural network(s) being trained. Thus, each training example can correspond to one operation of the artificial neural network(s) being trained. However, for each step, a batch of training samples is typically considered to accelerate the training process. Furthermore, numerous upgrades to gradient descent can also accelerate convergence or avoid local minima in the cost function landscape.

[0211] For each training multichannel audio signal, the training processor can perform a downmixing operation to generate a downmixed audio signal along with corresponding upmixing parameters and transient data. Therefore, the encoding process applied to the multichannel audio signal during normal operation can also be applied to the training multichannel audio signal to generate downmixing and upmixing parameter data.

[0212] Specifically, for stereo multichannel audio signals, the processor can be trained using a parameterized stereo scheme (e.g., according to a suitable normalization method). This encoding applies frequency- and time-dependent matrix operations (e.g., rotation operations) to the input stereo signal to generate a downmixed signal and a residual signal. For example, typically... Matrix multiplication / complex numerical multiplication is applied to the input stereo signal to, for example, substantially align one of the channels in the rotated signal to have the maximum signal value. This channel can be used as a single signal, and rotation is typically performed on a frame-by-frame basis. The rotation value can be stored as part of the upmixing parameter data (or the parameters that allow determining this rotation value can be included in the upmixing parameter data). Therefore, in the synthesis apparatus, the reverse rotation can be performed to reconstruct the stereo signal. The rotation of the stereo signal results in another stereo signal, whose channel is thus aligned with the maximum intensity. In parametric stereo encoders, the other channel is typically discarded to reduce the data rate. In conventional PS decoding, a decorrelation signal is typically generated at the decoder and used in the upmixing process. In current training methods, this second signal can be used as a residual signal from the downmixing because it can represent the information discarded in the encoder, and therefore it represents the ideal signal to be reconstructed in the decoder as part of the upmixing process.

[0213] Therefore, in some embodiments, the training processor can generate a training downmixed signal and a set of upmixing parameters and transient parameters (and possibly training residual signals) based on the training multi-channel audio signal. This training data can be fed into a decoder audio device including an artificial neural network 105 to generate an output multi-channel audio signal. A cost function is applied to determine the cost value of each training downmixed audio signal and / or a combination of training downmixed audio signals (e.g., determining the average cost value of the training set). The cost function may include various components.

[0214] Typically, the cost function will include at least one component that reflects how close the generated signal is to the reference signal, i.e., the so-called reconstruction error. In some embodiments, the cost function will include at least one component that reflects, from a perceptual perspective, how close the generated signal is to the reference signal.

[0215] Typically, the generated multi-channel audio signal can be compared with the original multi-channel audio signal input to the encoder's audio device, and the difference metric can be determined and used as the cost function. This process can be generated for all training sets to produce the total cost function.

[0216] It should be understood that many different methods can be used to determine the cost value that reflects the difference between signals. For example, a correlation method can be performed, where the cost value monotonically decreases as the correlation value increases. As another example, two signals can be subtracted from each other, and the power metric of the difference signal can be used as the cost value. It should be understood that many other methods can be used.

[0217] Therefore, in this example, the cost function generates a cost value that reflects the degree of matching between the generated multichannel audio signal and the corresponding original training multichannel audio signal.

[0218] Based on the cost value, the training processor can adapt the weights of the artificial neural network 105. For example, the backpropagation method can be used. Specifically, the training processor can adjust the weights of the artificial neural network 105 based on the cost value. For example, given the derivative of the weights with respect to the cost function (representing the slope), the weight values ​​are modified to move along the slope direction. For the simple / minimum case, it is possible to refer to the training of the perceptron (a single neuron) in the case of backpropagating a single data input.

[0219] This process can be iterated until the artificial neural network is considered to be trained. For example, training can be performed for a predetermined number of iterations. As another example, training can continue until the weight changes are less than a predetermined amount. Also very common is the implementation of validation stopping, where the network is tested again according to a validation metric and stops when the expected result is achieved.

[0220] In particular, audio devices (one or more) can be implemented in one or more appropriately programmed processors. Similarly, artificial neural networks can be implemented in one or more such appropriately programmed processors. Different functional blocks, especially artificial neural networks, can be implemented in separate processors and / or, for example, in the same processor. Examples of suitable processors are provided below.

[0221] Figure 17 This is a block diagram illustrating an example processor 1700 according to an embodiment of the present disclosure. Processor 1700 can be used to implement one or more processors that implement the means or elements thereof as described above (particularly including one or more artificial neural networks). Processor 1700 can be any suitable processor type, including but not limited to microprocessors, microcontrollers, digital signal processors (DSPs), field-programmable arrays (FPGAs) (wherein the FPGA has been programmed to form a processor), graphics processing units (GPUs), application-specific integrated circuits (ASICs) (wherein the ASIC has been designed to form a processor), or combinations thereof.

[0222] Processor 1700 may include one or more cores 1702. Core 1702 may include one or more arithmetic logic units (ALUs) 1704. In some embodiments, in addition to or in place of ALU 1704, core 1702 may include a floating-point logic unit (FPLU) 1706 and / or a digital signal processing unit (DSPU) 1708.

[0223] Processor 1700 may include one or more registers 1712 communicatively coupled to core 1702. Registers 1712 may be implemented using dedicated logic gates (e.g., flip-flops) and / or any memory technology. In some embodiments, registers 1712 may be implemented using static memory. Registers may provide data, instructions, and addresses to core 1702.

[0224] In some embodiments, processor 1700 may include a level-one or multi-level cache memory 1710 communicatively coupled to core 1702. Cache memory 1710 may provide computer-readable instructions to core 1702 for execution. Cache memory 1710 may provide data for core 1702 to process. In some embodiments, computer-readable instructions may be provided to cache memory 1710 from local memory (e.g., local memory attached to external bus 1716). Cache memory 1710 may be implemented using any suitable cache memory type, such as metal-oxide-semiconductor (MOS) memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology.

[0225] Processor 1700 may include controller 1714, which controls inputs to processor 1700 from other processors and / or components included in the system and / or outputs from processor 1700 to other processors and / or components included in the system. Controller 1714 may control data paths in ALU 1704, FPLU 1706, and / or DSPU 1708. Controller 1714 may be implemented as one or more state machines, data paths, and / or dedicated control logic. Gates of controller 1714 may be implemented as stand-alone gates, FPGAs, ASICs, or any other suitable technology.

[0226] Register 1712 and cache 1710 can communicate with controller 1714 and core 1702 via internal connections 1720A, 1720B, 1720C and 1720D. Internal connections can be implemented as buses, multiplexers, cross switches and / or any other suitable connection technology.

[0227] The inputs and outputs of processor 1700 may be provided via bus 1716, which may include one or more conductive lines. Bus 1716 may be communicatively coupled to one or more components of processor 1700, such as controller 1714, cache 1710, and / or register 1712. Bus 1716 may be coupled to one or more components of the system.

[0228] Bus 1716 may be coupled to one or more external memories. The external memory may include read-only memory (ROM) 1732. ROM 1732 may be a masked ROM, electrically programmable read-only memory (EPROM), or any other suitable technology. The external memory may include random access memory (RAM) 1733. RAM 1733 may be static RAM, battery-backed static RAM, dynamic RAM (DRAM), or any other suitable technology. The external memory may include electrically erasable programmable read-only memory (EEPROM) 1735. The external memory may include flash memory 1734. The external memory may include a magnetic storage device, such as a disk 1736. In some embodiments, the external memory may be included in the system.

[0229] This invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. Optionally, this invention can be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. Elements and components of embodiments of this invention can be implemented physically, functionally, and logically in any suitable manner. In practice, functionality can be implemented in a single unit, in multiple units, or as part of other functional units. Therefore, this invention can be implemented in a single unit or can be physically and functionally distributed among different units, circuits, and processors.

[0230] While the invention has been described in conjunction with some embodiments, it is not intended to be limited to the specific forms set forth herein. Rather, the scope of the invention is limited only by the appended claims. Furthermore, although it may appear that features have been described in conjunction with specific embodiments, those skilled in the art will recognize that various features of the described embodiments can be combined according to the invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.

[0231] Furthermore, although listed separately, multiple means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Additionally, while individual features may be included in different claims, these features may be advantageously combined, and inclusion in different claims does not imply that such a combination of features is impractical and / or disadvantageous. Moreover, including a feature in a claim of one class does not imply a limitation on that class, but rather indicates that the feature is equally applicable to other claim classes (as the case may be). Furthermore, the order of features in a claim does not imply any particular order in which the features must be performed, and in particular, the order of steps in a method claim does not imply that the steps must be performed in that order. Rather, the steps may be performed in any suitable order. Furthermore, singular references do not exclude plural. Therefore, references to “a,” “an,” “first,” “second,” etc., do not exclude plural. Reference numerals in the claims are provided as illustrative examples only and should not be construed in any way as limiting the scope of the claims.

Claims

1. An audio apparatus for generating and outputting multi-channel audio signals, the audio apparatus comprising: Receiver (101), which is arranged to receive audio data signals, said audio data signals including: The downmixed audio signal is the downmixed version of the first multi-channel audio signal; For the time-frequency segment of the downmixed audio signal, each set of upmixing parameters includes at least: The level difference parameter indicates the level difference between the channels of the first multi-channel audio signal; Correlation parameters, which indicate the coherence between the channels of the first multi-channel audio signal; and A phase difference parameter, which indicates the phase difference between the channels of the first multi-channel audio signal; and Neural network control data, which includes: At least one transient parameter indicating the transient characteristics of the first multi-channel audio signal, the transient parameter having a different time-frequency resolution than the set of upmixing parameters; An audio signal generator (103) is arranged to generate the output multi-channel audio signal by upmixing the downmixed audio signal according to the upmixing parameter set and the neural network control data. The audio signal generator (103) includes an artificial neural network (105) having an input node that receives the upmixing parameter set and an input node that receives the transient parameters.

2. The audio device according to claim 1, wherein, The audio signal generator (103) is arranged to generate the output multichannel audio signal by applying an upmixing coefficient to the downmixed audio signal and a decorrelation signal generated based on the downmixed audio signal, and the artificial neural network (105) is arranged to generate the upmixing coefficient.

3. The audio device according to any of the preceding claims, wherein, The audio signal generator (103) is arranged to generate a decorrelation signal based on the downmixed audio signal, and to generate at least one channel of the output multichannel audio signal by upmixing the downmixed audio signal and the decorrelation signal, and the artificial neural network (105) is arranged to control the generation of the decorrelation signal.

4. The audio device according to any of the preceding claims, wherein, The artificial neural network (105) includes inputs of segments of samples of the downmixed audio signal and outputs of segments of samples of the output multichannel audio signal.

5. The audio device according to any of the preceding claims, wherein, The neural network control data includes the inter-channel level difference for each of the multiple transients.

6. The audio device according to any of the preceding claims, wherein, The neural network control data includes timing parameters that indicate the timing of at least one transient state.

7. The audio device according to any of the preceding claims, wherein, The neural network control data does not include at least some transient inter-channel correlation or inter-channel phase difference data for the first multi-channel audio signal.

8. The audio device according to any of the preceding claims, wherein, The neural network control data has a lower frequency resolution than the upmixing parameters.

9. The audio device according to any of the preceding claims, wherein, The neural network control data includes data indicating the transient probability distribution characteristics of the first multi-channel audio signal.

10. An audio apparatus for generating audio data signals, the audio apparatus comprising: Receiver (201) receives a first multi-channel audio signal; A downmixer (203) is configured to: downmix the first multi-channel audio signal into a downmixed audio signal, and determine a set of upmixing parameters for time-frequency segments of the downmixed audio signal, each set of upmixing parameters including at least: The level difference parameter indicates the level difference between the channels of the multi-channel audio signal; Correlation parameters, which indicate the coherence between the channels of the multi-channel audio signal; and Phase difference parameter, which indicates the phase difference between the channels of the multi-channel audio signal; A transient detector (207) is arranged to determine at least one transient parameter indicating transient characteristics of the first multi-channel audio signal, the at least one transient parameter having a different time-frequency resolution than the set of upmixing parameters; and A generator (205) is arranged to generate the audio data signal to include the downmixed audio signal, the upmixed parameter set, and neural network control data including the at least one transient parameter.

11. The audio device according to claim 10, wherein, The transient detector (207) is arranged to detect a transient in response to detecting that a first level difference metric indicating the level difference between channels of the first multi-channel audio signal differs from a second level difference metric indicating the level difference between the channels by more than a threshold, the first level difference metric being determined for a time interval shorter than the second level difference; and to determine the at least one transient parameter to indicate the first level difference metric.

12. An audio system comprising: The audio apparatus for generating audio data signals according to claim 10 or 11, and the audio apparatus for generating output multi-channel audio signals based on the audio data signals according to any one of claims 1 to 9.

13. A method for generating and outputting multi-channel audio signals, the method comprising: Receive audio data signals, the audio data signals including: The downmixed audio signal is the downmixed version of the first multi-channel audio signal; For the time-frequency segment of the downmixed audio signal, each set of upmixing parameters includes at least: The level difference parameter indicates the level difference between the channels of the first multi-channel audio signal; Correlation parameters, which indicate the coherence between the channels of the first multi-channel audio signal; and A phase difference parameter, which indicates the phase difference between the channels of the first multi-channel audio signal; and Neural network control data, which includes: At least one transient parameter indicating the transient characteristics of the first multi-channel audio signal, the at least one transient parameter having a different time-frequency resolution than the set of upmixing parameters; The output multi-channel audio signal is generated by upmixing the downmixed audio signal according to the upmixing parameter set and the neural network control data. The output multi-channel audio signal is generated based on the output of an artificial neural network (207) having an input node that receives the upmixing parameter set and an input node that receives the transient parameters.

14. A method for generating an audio data signal, the method comprising: Receive the first multi-channel audio signal; The first multi-channel audio signal is downmixed into a downmixed audio signal, and a set of upmixing parameters for time-frequency segments of the downmixed audio signal is determined, wherein each set of upmixing parameters includes at least: The level difference parameter indicates the level difference between the channels of the multi-channel audio signal; Correlation parameters, which indicate the coherence between the channels of the multi-channel audio signal; and Phase difference parameter, which indicates the phase difference between the channels of the multi-channel audio signal; Determine at least one transient parameter indicating the transient nature of the first multi-channel audio signal, the at least one transient parameter having a time-frequency resolution different from the set of upmixing parameters; and The audio data signal is generated to include the downmixed audio signal, the upmixed parameter set, and neural network control data including the at least one transient parameter.

15. A computer program product comprising computer program code units adapted to perform all the steps of claim 13 or 14 when the program is run on a computer.