Multi-channel audio signal generation

JP2025529994A5Pending Publication Date: 2026-08-26KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025514543
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-13
Filing Date
2023-09-05
Publication Date
2026-08-26

AI Technical Summary

Technical Problem

Existing parametric stereo and multi-channel audio encoding/decoding techniques introduce distortions, variations, and artifacts, leading to degraded audio quality, high computational complexity, and inefficient data rates.

Method used

Employing a subband artificial neural network arrangement to process frequency subbands of a downmix audio signal, utilizing upmix parametric data to generate multi-channel audio signals, reducing complexity and resource usage.

Benefits of technology

Improves audio quality and reduces computational load while facilitating efficient generation and reconstruction of multi-channel audio signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The audio device comprises a receiver 101 for receiving a downmix audio signal for a multi-channel audio signal, upmix parametric data for upmixing the downmix audio signal, and a data signal comprising the upmix parametric data. A subband generator 103 generates frequency subband signals of the downmix audio signal, and a parameter generator 105 generates a set of upmix parameter values. The neural network arrangement 107, 401 comprises a plurality of subband artificial neural networks 107, 401 that receive the upmix parameter values ​​and samples of at least one frequency subband signal. The subband artificial neural networks 107, 401 generate subband samples for subbands of the frequency subband representation of the multi-channel audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the generation of multi-channel audio signals, and particularly, but not exclusively, to the decoding of a stereo signal from a downmix mono signal. [Background technology]

[0002] Spatial audio applications are growing in number and popularity, increasingly forming at least part of many audiovisual experiences. Indeed, new and improved spatial experiences and applications are constantly being developed, resulting in increased demands on audio processing and rendering.

[0003] For example, virtual reality (VR) and augmented reality (AR) have been gaining increasing interest in recent years, with numerous implementations and applications reaching the consumer market. Indeed, equipment is being developed for both providing experiences and capturing or recording suitable data for such applications. For example, relatively low-cost equipment is being developed to enable gaming consoles to provide full VR experiences. This trend is expected to continue and, indeed, accelerate as the VR and AR markets reach large scale within a short period of time. In the audio domain, a prominent field is exploring the reproduction and synthesis of realistic and natural spatial audio. The ideal goal is to produce a natural sound source such that a user cannot distinguish between the synthesized sound source and the original sound source.

[0004] Many research and development efforts have focused on providing efficient, high-quality audio encoding and decoding for spatial audio. Frequently used spatial audio representations are multi-channel audio representations, including stereo representations, and efficient encoding of such multi-channel audio has been developed based on downmixing a multi-channel audio signal using fewer channels into downmix channels. One of the main advances in low bitrate audio coding has been the use of parametric multi-channel coding, in which a downmix signal is generated together with parametric data that can be used to upmix the downmix signal to recreate the multi-channel audio signal.

[0005] In particular, instead of traditional mid-side or intensity coding, in parametric multi-channel audio coding, a multi-channel input signal is downmixed to fewer channels (e.g., from two to one) and multi-channel image (stereo) parameters are extracted. The downmix signal is then encoded using a more traditional audio coder (e.g., a mono audio encoder). The downmix bitstream is multiplexed with the coded multi-channel image parameter bitstream. This bitstream is then sent to a decoder, where the process is reversed. First, the downmix audio signal is decoded, and then the multi-channel audio signal guided by the coded multi-channel image / upmix parameters is reconstructed.

[0006] An example of stereo coding is described in E. Schuijers, W. Omen, B. den Brinker, and J. Breebaart, "Advances in Parametric Coding for High-Quality Audio," 114th AES Convention, Amsterdam, The Netherlands, 2003, Preprint 5852. In the described technique, the downmixed mono signal is parameterized by exploiting the natural separation of signals into three components (objects): transient, sinusoidal, and noise. A more detailed description is given in E. Schuijers, J. Breebaart, H. Purnhagen, and J. Engdegard, "Low Complexity Parametric Stereo Coding," 116th AES Convention, Berlin, Germany, 2004, Preprint 6073, which explains how parametric stereo with low (decoder) complexity is achieved when combined with spectral band replication (SBR).

[0007] In the described approach, decoding is based on the use of a so-called decorrelation process, which generates a decorrelated helper signal from the mono signal. In a stereo reconstruction process, both the mono signal and the decorrelated helper signal are used to generate an upmixed stereo signal based on the upmix parameters. In particular, the two signals are multiplied by a time- and frequency-dependent 2x2 matrix with coefficients determined from the upmix parameters to provide the output stereo signal. Summary of the Invention [Problem to be solved by the invention]

[0008] However, while parametric stereo (PS) and similar downmix encoding / decoding techniques were a leap from traditional stereo and multi-channel coding, they are not optimal in all scenarios. In particular, known encoding and decoding techniques tend to introduce certain distortions, variations, artifacts, etc. that result in differences between the (original) multi-channel audio signal input to the encoder and the reconstructed multi-channel audio signal at the decoder. Generally, audio quality is degraded and imperfect reconstruction of the multi-channels is performed. Furthermore, data rates are still higher than desired and / or processing complexity / resource usage is higher than preferred.

[0009] A further problem is the high complexity and computational load on the decoder side, which it would be desirable to reduce, especially for a given audio quality.

[0010] Therefore, improved techniques would be advantageous, particularly techniques that allow for increased flexibility, improved adaptability, improved performance, improved audio quality, trading off improved audio quality against data rate, reduced complexity and / or resource usage, reduced computational load, facilitated implementation, and / or improved audio experience.

[0011] SUMMARY OF THE INVENTION Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination. [Means for solving the problem]

[0012] According to one aspect of the present invention, an audio device for generating a multi-channel audio signal comprises: a receiver for receiving an audio data signal comprising a downmix audio signal for the multi-channel audio signal and upmix parametric data for upmixing the downmix audio signal; a subband generator for generating a set of frequency subband signals for subbands of the downmix audio signal; and an artificial neural network arrangement comprising a plurality of subband artificial neural networks, each subband artificial neural network of the plurality of subband artificial neural networks generating subband samples for a subband of the frequency subband representation of the multi-channel audio signal. an audio device comprising: a network configuration; a parameter generator for generating sets of upmix parameter values ​​for subbands of a frequency subband representation of a multi-channel audio signal from upmix parametric data; and a generator for generating a multi-channel audio signal from subband samples of the subbands of the multi-channel audio signal, wherein each subband artificial neural network comprises a set of nodes for receiving the set of upmix parameter values ​​and samples of at least one frequency subband signal of the set of frequency subband signals, the at least one frequency subband signal being for a subband for which the subband artificial neural network generates the subband samples of the multi-channel audio signal.

[0013] The present technique provides an improved audio experience in many embodiments. For many signals and scenarios, the technique provides improved generation / reconstruction of multi-channel audio signals, as well as improved perceived audio quality.

[0014] The present approach provides a particularly advantageous arrangement that in many embodiments and scenarios allows for facilitating and / or improving the possibilities of utilizing artificial neural networks in audio processing, including audio encoding and / or decoding in general, and allows for the advantageous employment of artificial neural networks in generating a multi-channel audio signal from a downmix audio signal.

[0015] The technique provides an efficient implementation and in many embodiments allows for reduced complexity and / or resource usage, and in many scenarios allows for a reduction in data rate for data representing a multi-channel audio signal using a downmix signal.

[0016] The subband samples may span a particular time and frequency range.

[0017] The upmix parametric data includes parameters (values) relating properties of the downmix signal to properties of the multi-channel audio signal. The upmix parametric data includes data indicative of relative properties between channels of a multi-channel audio signal. The upmix parametric data includes data indicative of differences in properties between channels of a multi-channel audio signal. The upmix parametric data includes data that is perceptually relevant to the synthesis of a multi-channel audio signal. Properties are, for example, differences in phase and / or magnitude and / or timing and / or correlation. In some embodiments and scenarios, the upmix parametric data expresses abstract properties that cannot be directly understood by humans / experts (but generally facilitate better reconstruction / lower data rates, etc.). The upmix parametric data includes data including at least one of inter-channel magnitude difference, inter-channel timing difference, inter-channel correlation and / or inter-channel phase difference for the channels of a multi-channel audio signal.

[0018] An artificial neural network is a trained artificial neural network.

[0019] The artificial neural network is a trained artificial neural network trained with training data including a training multi-channel audio signal, where the training employs a cost function that compares the training multi-channel audio signal to a multi-channel signal generated by the artificial neural network configuration. The artificial neural network is a trained artificial neural network trained with training data including training data representing a range of relevant audio sources, including recordings of music, video, motion pictures, telecommunications, etc.

[0020] The audio device is in particular an audio decoder device.

[0021] The subband generator generates one frequency subband signal for each subband of the downmix audio signal. Each frequency subband signal is represented by (subband) samples. Each subband artificial neural network has input nodes that receive the (subband) samples of the frequency subband signal (and possibly of two or more subbands of the downmix audio signal) generated by the subband generator.

[0022] Each subband artificial neural network generates an output for a subband of the multi-channel audio signal. Each subband artificial neural network generates subband samples representing the multi-channel audio signal in one subband in particular. A subband signal for one subband is generated by each subband artificial neural network. The subband signal generated for a subband includes the subband samples generated for that subband by the subband artificial neural network for that subband.

[0023] The subbands of the downmix audio signal generated by the subband generator are the same as the subbands of the multi-channel audio signal generated by the subband artificial neural network. However, in some embodiments, they are different. For example, multiple subbands of the downmix audio signal are fed to a single subband artificial neural network that generates subband samples for a single subband of the multi-channel audio signal.

[0024] Each subband artificial neural network comprises a set of nodes configured to receive samples for the subband; For the subbands, a subband artificial neural network generates subband samples for the multi-channel audio signal.

[0025] In many embodiments, the subbands have equal bandwidths.

[0026] The number of hidden layers in a subband artificial neural network is generally in the range of 2 to 100.

[0027] The number of input nodes in a subband artificial neural network is generally in the range of 16 to 1024.

[0028] The number of output nodes in a subband artificial neural network is generally in the range of 1 to 1024.

[0029] The number of nodes in the hidden layer of a subband artificial neural network is generally between 64 and 2048. 2 It is within the range of individuals.

[0030] The number of values ​​in one set of upmix parameters per subband is typically in the range of 1 to 6.

[0031] According to an optional feature of the invention, at least a first subband artificial neural network of the plurality of subband artificial neural networks comprises a node for receiving parameter values ​​of a set of upmix parameters for subbands other than the first subband of the subband artificial neural network.

[0032] This provides a particularly efficient implementation and / or improved performance.

[0033] According to an optional feature of the invention, at least some parameters of the sets of upmix parameter values ​​for the different subband artificial neural networks are the same.

[0034] This provides a particularly efficient implementation and / or improved performance.

[0035] In accordance with an optional feature of the invention, at least some parameters of the sets of upmix parameter values ​​for the different subband artificial neural networks are different.

[0036] This provides a particularly efficient implementation and / or improved performance.

[0037] In accordance with an optional feature of the invention, the plurality of subband artificial neural networks comprises, for at least one subband, separate artificial neural networks for different channels of the multi-channel audio signal.

[0038] This provides for a particularly efficient implementation and / or improved performance.In particular, separate subband artificial neural networks are provided for the left and right channel signals of a stereo multi-channel audio signal.

[0039] According to an optional feature of the invention, the parameter generator varies the resolution of the set of upmix parameters relative to the resolution of the upmix parametric data to match a resolution of the processing of the multiple subband artificial neural networks, the resolution of the processing of the multiple subband artificial neural networks being one of a frequency resolution of the subbands and a time resolution of a processing time interval for the multiple subband artificial neural networks.

[0040] In some embodiments, the upmix parametric data has a time resolution different from a processing time interval for at least one subband artificial neural network of the plurality of subband networks, and the parameter generator is configured to change the set time resolution of the upmix parameter values ​​to match the processing time interval for the at least one subband artificial neural network.

[0041] In some embodiments, the upmix parametric data has a frequency resolution that differs from the frequency resolution of the sub-bands of the downmix audio signal, and the parameter generator is configured to modify the frequency resolution of the set of upmix parameter values ​​to match the frequency resolution of the sub-bands of the downmix audio signal.

[0042] According to an optional feature of the invention, the parameter generator comprises at least one artificial neural network having a node for receiving parameter values ​​of the upmix parametric data and an output node for providing a set of upmix parameter values ​​for a first subband artificial neural network of the plurality of subband artificial neural networks.

[0043] This provides a particularly efficient implementation and / or improved performance.

[0044] The at least one artificial neural network comprises a plurality of sub-band artificial neural networks, and in some embodiments, one or more of the at least one artificial neural network is common to a plurality of sub-band artificial neural networks that generate samples of the multi-channel audio signal.

[0045] According to an optional feature of the invention, for at least a first subband, the plurality of subband artificial neural networks comprises at least two subband artificial neural networks that generate samples for different components of the subband signal for the first subband.

[0046] This provides a particularly efficient implementation and / or improved performance.

[0047] According to an optional feature of the invention, a plurality of subband artificial neural networks are trained with training data having training input audio signals comprising samples of the input multi-channel audio signal, and using a cost function including a component indicative of a difference between the training input audio signals and the multi-channel audio signals generated by the subband artificial neural networks.

[0048] This provides for particularly efficient implementation and / or improved performance, which in many embodiments provides for particularly efficient and high performance training.

[0049] According to an optional feature of the invention, a plurality of subband artificial neural networks are trained with training data having training input audio signals comprising samples of the input multi-channel audio signal, and using a cost function including a component indicative of a difference between upmix parameters for the training input audio signals and upmix parameters for the multi-channel audio signal generated by the subband artificial neural networks.

[0050] This provides for particularly efficient implementation and / or improved performance, which in many embodiments provides for particularly efficient and high performance training.

[0051] According to an optional feature of the invention, at least one subband artificial neural network of the plurality of subband artificial neural networks comprises: a first sub-artificial neural network having a node for receiving samples of frequency subband signals for the subband of the subband artificial neural network and an output node for supplying samples of a modified downmix audio signal; a second sub-artificial neural network having a node for receiving samples of frequency subband signals for the subband of the subband artificial neural network and an output node for supplying samples of an auxiliary audio signal; and a third sub-artificial neural network having a node for receiving samples of the modified downmix audio signal, a node for receiving samples of the auxiliary audio signal, and a node for receiving a set of upmix parameter values ​​for the subband of the subband artificial neural network, the third sub-artificial neural network further generating subband samples for the subband of the frequency subband representation of the multichannel audio signal.

[0052] This provides a particularly efficient implementation and / or improved performance.

[0053] According to an optional feature of the invention, the sets of upmix parameters have different numbers of parameters for at least two subbands.

[0054] This provides a particularly efficient implementation and / or improved performance.

[0055] According to an optional feature of the invention, the upmix parametric data provides parametric data for consecutive time intervals, and at least a first subband artificial neural network of the plurality of subband artificial neural networks comprises a node for receiving parameter values ​​of the set of upmix parameter values ​​for time intervals of the consecutive time intervals other than the time intervals for which the subband samples of the multi-channel audio signal are generated.

[0056] This provides a particularly efficient implementation and / or improved performance.

[0057] According to an optional feature of the invention, the plurality of subband artificial neural networks receives no input data other than subband samples of the downmix audio signal and parameter values ​​generated from the upmix parametric data.

[0058] This provides a particularly efficient implementation and / or improved performance.

[0059] According to an optional feature of the invention, there is provided a method for generating a multi-channel audio signal, the method comprising the steps of receiving an audio data signal comprising a downmix audio signal for the multi-channel audio signal and upmix parametric data for upmixing the downmix audio signal; generating a set of frequency subband signals for sub-bands of the downmix audio signal, wherein each sub-band artificial neural network of a plurality of sub-band artificial neural networks generates sub-band samples for a sub-band of the frequency sub-band representation of the multi-channel audio signal; generating a set of up-mix parameter values ​​for the sub-bands of the frequency sub-band representation of the multi-channel audio signal from the up-mix parametric data; and generating the multi-channel audio signal from the sub-band samples of the sub-bands of the multi-channel audio signal, each sub-band artificial neural network comprising a set of nodes receiving the set of up-mix parameter values ​​and samples of at least one frequency sub-band signal of the set of frequency sub-band signals, the at least one frequency sub-band signal being for a sub-band for which the sub-band artificial neural network generates sub-band samples of the multi-channel audio signal.

[0060] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter.

[0061] Embodiments of the present invention will now be described, by way of example only, with reference to the drawings in which: [Brief explanation of the drawings]

[0062] [Figure 1] FIG. 2 illustrates some elements of an example audio device, according to some embodiments of the present invention. [Figure 2]FIG. 1 illustrates an example of the structure of an artificial neural network. [Figure 3] FIG. 1 illustrates an example of a node of an artificial neural network. [Figure 4] FIG. 2 illustrates some elements of an example audio device, according to some embodiments of the present invention. [Figure 5] FIG. 2 illustrates some elements of an example audio device, according to some embodiments of the present invention. [Figure 6] FIG. 1 illustrates some elements of an example apparatus for training an artificial neural network for an audio device, according to some embodiments of the present invention. [Figure 7] 2A-2C illustrate some elements of possible configurations of a processor for implementing elements of an audio device according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0063] FIG. 1 illustrates some elements of an audio device according to some embodiments of the present invention.

[0064] The audio device comprises a receiver 101 configured to receive a data signal / bitstream comprising a downmix audio signal that is a downmix of a multi-channel audio signal. The following description focuses on the case where the multi-channel audio signal is a stereo signal and the downmix signal is a mono signal, but it will be understood that the techniques and principles described are equally applicable to multi-channel audio signals having more than two channels, and to downmix signals having more than two channels (although fewer channels than the multi-channel audio signal).

[0065] Furthermore, the received data signal includes upmix parametric data for upmixing the downmix audio signal. The upmix parametric data is in particular a set of upmix parameters that indicate a relationship between signals of different audio channels of a multi-channel audio signal (in particular a stereo signal) and / or between the downmix signal and the audio channels of the multi-channel audio signal. Typically, the upmix parameters indicate a measure of similarity, such as a time difference, a phase difference, a level / intensity difference, and / or a correlation. Typically, the upmix parameters are provided per time and per frequency (time-frequency tile). For example, new parameters for a set of subbands are provided periodically. The parameters include in particular an inter-channel phase difference (IPD) parameter, an overall phase difference (OPD) parameter, an inter-channel correlation (ICC) parameter, and a channel phase difference (CPD) parameter, which are known from parametric stereo coding (as well as from higher channel coding).

[0066] Typically, the downmix audio signal is encoded and the receiver 101 includes a decoder for decoding the downmix audio signal, i.e., in the particular example, a mono signal. It will be understood that a decoder is not required if the received downmix audio signal is not encoded and that the decoder is considered to be an integral part of the receiver. Similarly, the receiver 101 includes functionality for extracting and decoding data representing the upmix parameters.

[0067] Traditionally, decoding of a signal, such as a PS-encoded stereo signal, is based on generating a decorrelated signal from a downmix audio signal (especially a mono signal) and then applying a (time- and frequency-dependent) 2×2 matrix multiplication between samples of the downmix audio signal and the decorrelated signal, resulting in an output multi-channel audio signal. The coefficients of the 2×2 matrix are determined from upmix parameters of the upmix parametric data. However, while such an approach is suitable for many applications, it is not ideal in all situations and tends to have suboptimal performance in some scenarios. The approach of FIG. 1 uses a fundamentally different approach that offers improved performance and / or easier implementation in many embodiments and scenarios.

[0068] 1, the receiver 101 is coupled to a subband generator 103 configured to generate a plurality of frequency subband signals for subbands of the downmix audio signal. Thus, a subband representation of the downmix audio signal is generated, the subband generator 103 generating subband samples for different frequency subbands, and therefore it generates a plurality of subband samples for the different subbands.

[0069] In particular, the subband generator 103 includes a filter bank configured to generate frequency subband representations of the downmix audio signal. The filter bank may be a quadrature mirror filter (QMF) bank and may be implemented, for example, by a Fast Fourier Transform (FFT), although it will be understood that many other filter banks and techniques for splitting an audio signal into multiple subband signals are known and may be used. The filter bank is in particular a complex-valued pseudo-QMF bank, yielding, for example, 32 or 64 complex-valued subband signals.

[0070] In many embodiments, the filter bank 501 is configured to generate a set of subband signals for subbands having equal bandwidths. In other embodiments, the filter bank 401 is configured to generate subband signals using subbands having different bandwidths. For example, higher frequency subbands have higher bandwidths than lower frequency subbands. Also, subbands are grouped together to form higher bandwidth subbands.

[0071] Generally, the subbands have bandwidths in the range of 10 Hz to 10,000 Hz.

[0072] The audio device further comprises a parameter generator 105 configured to generate a set of upmix parameters for subbands of the downmix audio signal from the received upmix parametric data. In some embodiments, the parameter generator 105 may simply pass on the received upmix parameters without modification, but may also select and distribute appropriate upmix parameters to other functional units. In other embodiments, the parameter generator 105 processes the received upmix parameter values ​​to generate new parameter values, for example by interpolation and / or upsampling.

[0073] The audio device further comprises an artificial neural network arrangement comprising a plurality of subband artificial neural networks 107, only one of which is shown in FIG. 1 for clarity. In many embodiments, the artificial neural network arrangement comprises one subband artificial neural network 107 for each subband generated by the subband generator 103. Each of the subband artificial neural networks 107 is configured to generate subband samples for that subband of the frequency subband representation of the multi-channel audio signal. Each subband artificial neural network 107 comprises a set of output nodes that generate samples for that subband of the multi-channel audio signal to be reconstructed. The subband artificial neural network 107 has nodes configured to receive subband samples of the downmix audio signal; in particular, subband samples of one or more subbands corresponding to the frequencies of the subbands of the multi-channel audio signal for which the subband artificial neural network 107 is generating output samples are supplied to the input nodes of the subband artificial neural network 107. Furthermore, the subband artificial neural network 107 comprises a node that receives a set of upmix parameter values ​​for the subband from the parameter generator 105. In many embodiments, each subband artificial neural network 107 receives a set of upmix parameters received in the upmix parametric data, the upmix parameters being given for the subbands of the particular subband artificial neural network 107.

[0074] In many embodiments, the subbands generated by the subband generator 103 and the subbands for which the subband artificial neural network 107 generates samples are the same. There is a direct correspondence between the subbands of the downmix audio signal and the subbands of the multi-channel audio signal. In particular, there is one subband artificial neural network 107 for each subband generated by the subband generator 103, each of which generates subband samples of the multi-channel audio signal for the same subband. In some embodiments, several subbands of the downmix audio signal are combined, for example, to be processed by the same subband artificial neural network 107 (which can be considered equivalent to one subband having the combined bandwidth of the combined subbands). The subband artificial neural network 107 then generates subband samples for the combined subbands.

[0075] As will be explained in more detail later, the subband artificial neural network 107 is trained to generate subband samples that reconstruct a multi-channel audio signal from the downmix audio signal and the upmix parameters. In this approach, the downmix audio signal is therefore divided into subbands, which are processed by the trained subband artificial neural network 107, which directly generates subband samples of the multi-channel audio signal.

[0076] The subband artificial neural network 107 is coupled to a signal generator 109 which generates a multi-channel audio signal from sub-band samples of the sub-bands of the multi-channel audio signal.

[0077] The subband samples from the subband artificial neural network are provided to a signal generator 109, which proceeds to generate a reconstructed multi-channel audio signal. For example, in some embodiments where a subband representation of the multi-channel audio signal is desired (e.g., by subsequent processing that is also subband-based), the signal generator 109 simply outputs the subband samples from the subband artificial neural network, possibly according to a particular structure or format. In many embodiments, the signal generator 109 comprises functionality for converting the subband representation of the reconstructed multi-channel audio signal into a time-domain representation. The signal generator 109 comprises, in particular, a synthesis filterbank that performs the inverse operation of the subband generator 103, specifically the filterbank of the subband generator 103, thereby converting the subband representation into a time-domain representation of the multi-channel audio signal.

[0078] The generator is particularly configured to generate the frequency / subband domain representation of the multi-channel audio signal by processing the frequency or subband domain representation of the downmix audio signal and the frequency / subband domain representation of the auxiliary audio signal. The processing of the generator 105 is therefore subband processing, e.g. a matrix multiplication performed in each subband on the subband samples of the downmix audio signal and the auxiliary audio signal generated by the corresponding subband artificial neural networks.

[0079] The resulting subband / frequency domain representation is then either used directly or converted to a time domain representation using, for example, a suitable synthesis filter bank applied specifically with a separate synthesis filter for each channel.

[0080] The artificial neural network used in the described functions is a network of nodes organized into layers, each node having a node value. Figure 2 shows an example of a section of an artificial neural network.

[0081] The node value for a given node is calculated to include contributions from some, or often all, nodes in previous layers of the artificial neural network. In particular, the node value for a node is calculated as a weighted sum of the node values ​​of all node outputs in the previous layer. Typically, a bias is added, and the result is applied to an activation function. Activation functions generally provide an essential part of each neuron by providing nonlinearity. Such nonlinearity and activation functions provide significant benefits in the learning and adaptation process of neural networks. Thus, node values ​​are generated as a function of the node values ​​in the previous layer.

[0082] The artificial neural network comprises, inter alia, an input layer 201 comprising a plurality of nodes that receive input data values ​​for the artificial neural network. Thus, the node values ​​for the nodes of the input layer generally become input data values ​​to the artificial neural network directly and are therefore not calculated from other node values.

[0083] An artificial neural network may further comprise zero, one, or more hidden layers 203 or processing layers. For each such layer, node values ​​are generally generated as a function of the node values ​​of the nodes in the previous layer, in particular by applying a weighted combination and an additional bias followed by an activation function.

[0084] In particular, as shown in Figure 3, each node, sometimes called a neuron, receives input values ​​(from nodes in previous layers) and then calculates the node value as a function of these values. Often, this involves first generating the value as a linear combination of the input values, where each input value is represented by a weight

number

[0085] An activation function is then applied to the resulting combination. For example, the node value l is l=f(k) where the function can be, for example, a sigmoid, Tanh, or Rectified Linear Unit (ReLU) function (as described in Xavier Glorot, Antoine Bordes, Yoshua Bengio Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, PMLR 15:315-323, 2011), i.e., f(k) = ReLU(k) = max(0,k) is.

[0086] Other frequently used functions include the sigmoid function or the tanh function. In many embodiments, node outputs or values ​​are calculated using multiple functions. For example, f(k)=ReLU(k)=σ(k) Both ReLU and sigmoid functions are combined using activation functions such as

[0087] Such operations are performed by each node of the artificial neural network (generally except for the input nodes).

[0088] The artificial neural network further comprises an output layer 205 that provides output from the artificial neural network, i.e., the output data of the artificial neural network are the node values ​​of the output layer. For hidden / processing layers, the output node values ​​are generated by functions of the node values ​​of the previous layer. However, in contrast to the hidden / processing layers, where the node values ​​are generally inaccessible or not further used, the node values ​​of the output layer are accessible and provide the results of the operation of the artificial neural network.

[0089] Several different network structures and toolboxes for artificial neural networks have been developed, and in many embodiments, artificial neural networks are based on adapting and customizing such networks. One example of a network architecture suitable for the above-mentioned applications is WaveNet by van den Oord et al., described in Oord, Aaron van den, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. "Wavenet: A generative model for raw audio." arXiv preprint arXiv:1609.03499 (2016).

[0090] WaveNet is an architecture used for synthesis of time-domain signals using dilated causal convolution, and has been successfully applied to audio signals. In WaveNet, the following activation function is used:

number

number

[0091] Artificial neural networks are sometimes further configured to include additional contributions that allow the artificial neural network to be dynamically adapted or customized for particular desired properties or characteristics of the generated output. For example, a set of values ​​is provided to adapt the artificial neural network. These values ​​are included by providing contributions to some nodes of the artificial neural network. These nodes are specifically input nodes, but generally nodes in hidden or processing layers. Such adaptation values ​​are weighted and added, for example, as contributions to a weighted sum / correlation value for a given node. For example, in the case of WaveNet, such adaptation values ​​are included in an activation function. For example, the output of the activation function is

number

[0092] The above description relates to a neural network approach that is suitable for many embodiments and implementations. However, it will be understood that many other types and structures of neural networks may be used. Indeed, many different approaches for generating neural networks have been and are being developed, including neural networks that use complex structures and processes different from those described above. The approach is not limited to any particular neural network approach, and any suitable approach may be used without detracting from the invention.

[0093] 1, a single subband artificial neural network 107 is used to generate subband samples for multiple channels of a multi-channel audio signal in a given subband, and in particular, a single subband artificial neural network 107 is used to generate subband samples for both channels of a stereo signal. Thus, the subband artificial neural network 107 generates output samples for both the left and right channels.

[0094] However, in many embodiments, a separate sub-band artificial neural network 107 is provided for each channel of the multi-channel audio signal, and therefore a parallel structure and configuration of sub-band artificial neural networks 107 is provided for each channel.

[0095] 4 shows an example of an audio device corresponding to that of FIG. 1, but with separate subband artificial neural networks 107, 401 for the left and right channels, respectively. In this example, the subband artificial neural network 107 is configured and trained to generate samples of the left signal of a stereo signal. Additionally, a second subband artificial neural network 401 is configured and trained to generate samples of the right signal of the stereo signal. Both the first and second subband artificial neural networks 107, 401 have input nodes that receive subband samples from the subband generator 103. Similarly, both the first and second subband artificial neural networks 107, 401 receive upmix parameters from the parameter generator 105, although in some embodiments, they receive different upmix parameters. For example, one or more upmix parameters may specifically indicate properties of one channel signal relative to the downmix audio signal, and such parameters may be provided only to the subband artificial neural networks 107, 401 for the channel to which the upmix parameters relate.

[0096] In such an embodiment, separate signal generators are provided for and applied to the respective sub-band artificial neural networks 107, 401. In particular, the signal generator 109 receives output samples from the first sub-band artificial neural network 107, and the second separate signal generator 403 receives sub-band samples from the second sub-band artificial neural network 401.

[0097] In this approach, the left and right channels of a multi-channel audio signal are obtained directly from two pre-trained artificial neural networks that take the downmix signal and the decoded upmix parameters as input. This approach is performed on a subband basis, such that reconstruction and synthesis of the multi-channel audio signal is achieved by a combination of multiple subband artificial neural networks, each of which is provided for a different subband, interacting to provide a direct generation of the multi-channel audio signal. The artificial neural network configuration does not require or involve any decorrelation signals or rotation / matrix multiplications to be performed (although the artificial neural networks may potentially produce operations corresponding to broader (and less clearly separable) variants of the operation). This approach has been found to provide significant improvements in the reconstruction of multi-channel audio signals in many scenarios and embodiments. Furthermore, this approach facilitates implementation and often reduces computational burden. For example, the use of a subband approach often reduces complexity and resource usage, since individual subband artificial neural networks are typically significantly smaller than a single artificial neural network required to generate samples for all frequencies of a multi-channel audio signal, for example. Although more artificial neural networks are used for subband processing, the size reduction that can be achieved is often much greater than simple linear scaling, and therefore an overall complexity reduction can be achieved that requires significantly less computation and operations.

[0098] The subband artificial neural network 107 receives, in particular, the subband samples of the downmix audio signal as well as parameter values ​​determined from the received upmix parametric data from the subband generator 103, but is generally not provided with any other input data. Thus, the present approach does not require any additional information, operations or data other than that of the coded signal representing the multi-channel audio signal, and in particular does not require any data other than the subband samples and the upmix parameters.

[0099] In this configuration, each of the subband artificial neural networks receives subband samples for its subband, and further, all of the subband artificial neural networks are configured to receive a set of upmix parameter values.

[0100] Each of the subband artificial neural networks generates subband samples for a subset of the subbands of the frequency subband representation of the multi-channel audio signal, and typically generates subband samples (only) for the subbands for which it receives input samples from the subband generator 103.

[0101] In many embodiments, the apparatus includes an artificial neural network for each subband of the frequency subband representation of the downmix audio signal generated by the subband generator 103. Thus, in many embodiments, the output samples for each subband of the subband generator 103 are fed to input nodes of one subband artificial neural network, which then generates subband samples of the multi-channel audio signal for that subband. In many embodiments, the subband processing is therefore completely separate for each subband.

[0102] In this example, the generation of the multi-channel audio signal is therefore performed subband by subband, with a separate and individual artificial neural network in each subband, which is trained to provide output samples for the subbands to which it is supplied with input subband samples.

[0103] Such an approach has been found to provide highly advantageous generation of multi-channel audio signals, enabling extremely high-quality reconstruction of the multi-channel audio signal. Furthermore, such an approach allows for highly efficient operation with significantly reduced complexity and / or generally significantly reduced computational resource requirements. Subband artificial neural networks tend to be significantly smaller than a single complete artificial neural network required for the generation of the entire signal. Generally, far fewer nodes and, in some cases, fewer layers are required for processing, resulting in a significant reduction in the number of operations and calculations required to perform the artificial neural network function. While more artificial neural networks are required to cover all subbands, smaller artificial neural networks generally significantly reduce the total number of operations required, and therefore the overall computational resource requirements. Furthermore, in many scenarios, this allows for a more efficient training process.

[0104] The subband structure thus provides a computationally efficient approach for enabling an artificial neural network to be implemented to assist in the reconstruction of a multi-channel audio signal that has been coded as a downmix audio signal with upmix parametric data. The described systems and techniques enable high-quality multi-channel audio signals to be reconstructed, and in general, significantly improved audio quality can be achieved compared to conventional techniques. Furthermore, a computationally efficient decoding process can be achieved. The subband and artificial neural network-based techniques are also compatible with other processes that use subband processing.

[0105] In some embodiments, subband processing is more flexible than strict subband-by-subband processing. For example, in some embodiments, each subband artificial neural network receives not only subband samples from the subband itself, but also, possibly, subband samples for one or more other subbands. For example, a subband artificial neural network for one subband may, in some embodiments, also receive samples of the downmix audio signal from one or two neighboring / adjacent subbands. As another example, in some embodiments, one or more of the subband artificial neural networks may also receive input samples from one or more subbands that contain harmonics (or subharmonics) of the frequencies of the subband. For example, a subband around a 500 Hz center frequency may also receive frequencies from a subband around a 1000 Hz center frequency. Such additional subbands, having a specific relationship with the subbands of the subband artificial neural network, provide additional information that allows for the generation of improved subband artificial neural networks for some audio signals.

[0106] In some embodiments, all sub-band artificial neural networks have the same properties and dimensions. In particular, in many embodiments, all sub-band artificial neural networks have the same number of input and output nodes, and possibly the same internal structure. Such an approach is used, for example, in embodiments in which all sub-bands have the same bandwidth.

[0107] In some embodiments, however, the sub-band artificial neural networks comprise non-equivalent neural networks. In particular, in some embodiments, the number of input nodes for the sub-band artificial neural networks is different for at least two of the artificial neural networks. Thus, in some embodiments, the number of input samples included in determining the output samples is different for different sub-bands and sub-band artificial neural networks.

[0108] In some embodiments, the number of samples / input nodes is greater in some lower frequency subbands than in some higher frequency bands. In fact, the number of samples / input nodes decreases monotonically with increasing frequency. Lower frequency subband artificial neural networks are therefore larger than higher frequency subband artificial neural networks and consider more input samples than higher frequency subband artificial neural networks. Such an approach can be combined with subbands having different bandwidths, for example, when lower frequency subbands have a higher bandwidth than higher frequency bands.

[0109] Such an approach offers, in many scenarios, an improved trade-off between achievable audio quality and computational complexity and resource usage, allowing for closer adaptation of the system to reflect the general characteristics of audio, thereby enabling more efficient processing.

[0110] In some embodiments, the subband artificial neural network is employed only for a subset of the subbands of the downmix audio signal and / or the multi-channel audio signal. For other subbands, other approaches are applied, in particular for one or more subbands, a conventional approach is applied to generate a decorrelated signal that is mixed with the mono downmix audio signal to generate a stereo signal. Therefore, the described subband artificial neural network approach is applied only for some subbands.

[0111] In some embodiments, the number of hidden layers is greater for some lower frequency bands than for some higher frequency bands, and the number of hidden layers decreases monotonically as frequency increases. Such an approach offers an improved trade-off between achievable audio quality and computational complexity and resource usage in many scenarios.

[0112] In some embodiments, at least some of the sets of upmix parameter values ​​are common for multiple subbands of the frequency subband representation of the downmix audio signal.

[0113] In some embodiments, all of the sub-band artificial neural networks are provided with the same set of upmix parameter values. In some embodiments, only some of the sub-band artificial neural networks are provided with the same set of upmix parameter values. In particular, in some embodiments, at least one control data value of the set of control data values ​​is processed by at least two combined sub-band artificial neural networks.

[0114] Using the same set of upmix parameter values ​​provides improved efficiency and performance in many embodiments, often reducing the complexity and resource usage in generating the set of upmix parameter values. Furthermore, in many scenarios, it provides improved performance, as all available information provided by the parameter values ​​is considered by each subband artificial neural network, and thus improved adaptation of the subband artificial neural networks is achieved.

[0115] However, in many embodiments, different subband artificial neural networks are provided with different sets of upmix parameter values, and in particular, in some embodiments, at least one parameter value of the set of upmix parameter values ​​is not processed by at least one subband artificial neural network.

[0116] For example, the parameter generator 105 generates a set of upmix parameter values, and different subsets of these are provided to different subband artificial neural networks. In other embodiments, some parameter sets are also generated to include some parameter values ​​that are provided manually or generated, for example, by analysis of the downmix audio signal. For example, harmonics or peaks are detected in the downmix audio signal. Such data is then applied, for example, to only some of the synthesis subband artificial neural networks. For example, detected peaks or harmonics are only shown to the synthesis subband artificial neural networks of the subbands in which they were detected. In some embodiments, such properties and features are generated at the encoder side and provided to the audio device as part of the data signal. Such received features are also provided to the artificial neural networks, i.e., the subband artificial neural networks include input nodes for receiving values ​​representing such properties.

[0117] In some embodiments, the encoder alternatively or additionally generates upmix parametric data, e.g., in the form of data representing or describing properties of the downmix audio signal, e.g., with respect to properties of one or more channels of the multi-channel audio signal. Indeed, in some embodiments, the encoder comprises an artificial neural network that generates, for the downmix audio signal, a set of parameter values ​​that provide a latent representation that is particularly suitable for upmixing the downmix audio signal to reconstruct the multi-channel audio signal, the upmixing being performed by an artificial neural network configuration as described herein. In such a case, the encoder artificial neural network that generates the latent representations / parameter values ​​is trained together with the subband artificial neural network 107.

[0118] In many embodiments, different subband artificial neural networks are therefore given different sets of upmix parameter values, which often involves some parameter values ​​being the same and some parameter values ​​being different for the different subband artificial neural networks.

[0119] As mentioned above, the parameter generator 105, in some embodiments, generates sets of upmix parameters for the different subband artificial neural networks 107 by selecting appropriate parameters from the received upmix parametric data and feeds these directly to the appropriate subband artificial neural networks 107.

[0120] For example, in many embodiments, the upmix parametric data includes upmix parameters that are frequency- and time-dependent. For example, IPD parameters, OPD parameters, ICC parameters, and CPD parameters are provided for separate time-frequency tiles. In such cases, the upmix parameters provided in the upmix parametric data for the frequency subbands of one subband artificial neural network 107 for a given time interval are then compiled into a set of upmix parameters that are supplied to the subband artificial neural network 107 when processing the given time interval. Thus, when a given subband artificial neural network 107 is generating subband samples for a given time interval, the parameter generator 105 generates a set of upmix parameters for that time interval that includes the upmix parameters provided for that subband. Furthermore, the subband artificial neural network 107 receives subband samples of the downmix audio signal, thereby enabling the subband artificial neural network 107 to process this input data to generate subband samples for the reconstructed multi-channel audio signal.

[0121] In some embodiments, each set of upmix parameters includes parameter values ​​for only the subbands for which the subband artificial neural network 107 generates samples, and therefore only the parameter values ​​for the subbands of a particular subband artificial neural network 107 are provided to that subband artificial neural network 107. However, in some embodiments, the set of upmix parameters for one subband artificial neural network 107 includes parameter values ​​for subbands other than the subbands for which the subband artificial neural network 107 generates samples. Thus, in some embodiments, the subband artificial neural network 107 comprises a node for receiving parameter values ​​for the subbands other than that subband of that subband artificial neural network.

[0122] In some embodiments, the parameter generator 105 generates a set of upmix parameter values ​​for a given subband based on the received upmix parametric data for that subband. However, one or more of the subband artificial neural networks also includes parameter values ​​for one or more other subbands in addition to the set generated for that subband of the subband artificial neural network, i.e., the input to a subband artificial neural network has input nodes that receive subband samples for subbands for which that subband artificial neural network does not generate any samples.

[0123] As a specific example, in many embodiments, each subband artificial neural network receives as input not only parameter values ​​from the subband from which it generates a set of samples, but also parameter values ​​from, for example, adjacent subbands.

[0124] Such an approach often allows for generating an improved set of upmix parameter values, which often leads to an improved audio quality. In particular, it has been found that considering surrounding subbands allows the set of upmix parameter values ​​to better reflect the temporal resolution of the downmix audio signal / multi-channel audio signal. It has been found that such an approach in particular allows for a better representation of temporal peakedness.

[0125] In some embodiments, one or more of the subband artificial neural networks are also supplied with subband samples from outside the time interval in which the subband artificial neural network generates samples of the multi-channel audio signal, and the subband artificial neural network includes an input node that receives subband samples from outside the current time interval.

[0126] In particular, the audio device processing is performed frame-by-frame, where a time interval / frame of the received downmix audio signal is processed to generate output samples for the multi-channel audio signal for that time interval / frame. Thus, for each frame, the subband generator 103 generates subband samples, the parameter generator 105 generates parameter values ​​for that frame, and these subband samples and parameter values ​​are fed to a subband artificial neural network which generates subband samples for the multi-channel audio signal for that frame / time interval of the multi-channel audio signal.

[0127] Thus, in particular, each subband artificial neural network operates in a block manner, each operation in which a set of output samples is generated from a set of input samples corresponding to a time interval of the downmix audio signal / multi-channel audio signal in which an output sample of the multi-channel audio signal is generated.

[0128] In some embodiments, one or more of the subband artificial neural networks, in addition to the parameter values ​​provided for that subband, also receive parameter values ​​for another time interval, such as typically from one or more adjacent time intervals. For example, in some embodiments, one or more of the subband artificial neural networks also includes parameter values ​​for the previous and next time intervals.

[0129] In such an example, the upmix parametric data thus provides parameters for a number of different consecutive time intervals. For a given time interval, the subband artificial neural networks 107 are provided with parameter values ​​for the time interval of the multi-channel audio signal for which the subband samples are generated. However, in addition, in some embodiments, one or more of the subband artificial neural networks 107 also comprise nodes for receiving parameter values ​​for time intervals of that consecutive time interval other than the time interval for which the subband samples of the multi-channel audio signal are (currently) generated.

[0130] In some embodiments, the subband samples provided to a subband artificial neural network 107 are only for the subbands for which the subband artificial neural network 107 generates samples, and thus the subband samples for the subbands of an individual subband artificial neural network 107 are provided only to that subband artificial neural network 107. However, in some embodiments, the subband samples for one subband artificial neural network 107 include subband samples for subbands other than the subband for which the subband artificial neural network 107 generates samples. Thus, in some embodiments, the subband artificial neural network 107 comprises a node for receiving subband samples for subbands other than that subband of that subband artificial neural network.

[0131] As a specific example, in many embodiments, each subband artificial neural network receives as input not only subband samples from the subband for which it generates a set of upmix parameter values, but also subband samples from, for example, adjacent subbands.

[0132] Such an approach often leads to an improvement in audio quality. In particular, it has been found that considering surrounding subbands allows the generated subband samples to better reflect the temporal resolution of the downmix audio signal / multi-channel audio signal. In particular, it has been found that such an approach allows a better representation of temporal peakedness.

[0133] Each subband artificial neural network operates on a processing time interval, and each operation in which a set of output samples is generated from a set of input samples corresponds to a time interval of the downmix audio signal / multi-channel audio signal in which an output sample of the multi-channel audio signal is generated.

[0134] In some embodiments, one or more of the subband artificial neural networks also receive subband samples for another time interval, such as typically from one or more adjacent time intervals, in addition to the subband samples provided for that time interval. For example, in some embodiments, one or more of the subband artificial neural networks also includes subband samples for previous and next time intervals.

[0135] In general, the upmix parameters received in the upmix parametric data have a time and frequency resolution that differs from the subband domain downmix audio signal and from the subbands and processing time interval of the subband artificial neural network 107. In many embodiments, the parameter generator 105 is configured to adapt the received upmix parameters to generate a set of upmix parameter values ​​that match the processing time and frequency resolution of the subband artificial neural network 107.

[0136] In many embodiments, the parameter generator 105 is therefore configured to change the resolution of the set of upmix parameters relative to the resolution of the upmix parametric data to match the resolution of the processing of the multiple subband artificial neural networks 107. The change of resolution is in the frequency and / or time domain and is performed to match the upmix parametric data to the frequency resolution of the subbands for the multiple subband networks and / or to the time resolution for the processing time interval.

[0137] In some embodiments, the change in resolution is effectively a resampling of the received parameter values ​​to match the time and frequency resolution of the subband processing. For example, in some embodiments, linear interpolation is applied to each parameter to generate sample values ​​of the parameter for a time and frequency interval corresponding to the subband processing. For example, if the upmix parametric data includes parameter values ​​for two frequencies corresponding to two adjacent subbands that are larger than the subbands of the subband processing, the parameter values ​​for the processing subbands are found by simple interpolation between the parameter values ​​included in the upmix parametric data. Similarly, if the parameter values ​​of the upmix parametric data are for a time interval larger than the processing time interval, interpolation may be applied to generate higher resolution parameter values ​​for the set of upmix parameters. It will be understood that many different techniques for resampling (both to increase and decrease resolution) are known to those skilled in the art, and any suitable technique may be applied.

[0138] In some embodiments, the parameter generator 105 advantageously comprises one or more artificial neural networks. In some embodiments, the parameter generator 105 comprises a single artificial neural network that generates sets of upmix parameter values ​​for all sub-band artificial neural networks 107. However, in many embodiments, the parameter generator 105 comprises artificial neural networks for multiple, and typically all, sub-band artificial neural networks 107. Thus, in some embodiments, the parameter generator 105 includes one artificial neural network for each sub-band artificial neural network 107.

[0139] The use of a trained artificial neural network to generate the set of upmix parameters provides improved operation and performance in many scenarios, and in particular, the use of a subband artificial neural network to generate the set of upmix parameters provides improved performance while maintaining low complexity and computational resource usage.

[0140] In some embodiments, the set of upmix parameters includes the same number of parameters for each subband artificial neural network 107, and each subband artificial neural network 107 has the same number of nodes that receive contributions from the upmix parameters. For example, in some scenarios, each parameter set includes one set of complementary parameters included in the upmix parametric data for a subband of the subband artificial neural network 107. For example, for each subband and subband artificial neural network 107, a set of upmix parameters is provided that includes IID parameters, IPD parameters, and ICC parameters. Thus, in such embodiments, each subband artificial neural network 107 receives one IID parameter, one IPD parameter, and one ICC parameter that reflect the upmixing for that subband.

[0141] However, in other embodiments, the set of upmix parameters has a different number of parameters for at least two subbands. For example, in some embodiments, for some subbands only one or two of the IID, IPD and ICC parameters are included, while for other subbands all parameters are provided. This may reflect, for example, that some parameters are more important for some frequency ranges than others.

[0142] As another example, in some embodiments, some subbands have different sizes, and the upmix parametric data includes more upmix parameters for some subbands than for others. The parameter generator 105, for example, generates more parameters for some subbands than for others (e.g., for different frequency subranges within each subband). Subband artificial neural networks 107 that cover larger bandwidths and for which more upmix parameters are generated are therefore configured to have more nodes that receive input from the upmix parameters than other subband artificial neural networks 107.

[0143] As another example, in many embodiments, the update rate for the upmix parameters is different for different frequency ranges, and therefore there are (or are generated by the parameter generator 105) more parameter values ​​for some subbands than for other subbands. For example, in some embodiments, one set of upmix parameters is provided for some subbands, while multiple sets are provided for other subbands for different times within the processing interval. The subband artificial neural network 107 for these subbands is configured with different nodes for such parameter values ​​to be included in determining the subband samples of the multi-channel audio signal.

[0144] Such techniques generally allow for improved multi-channel audio signal reconstruction and allow processing to be more accurately adapted to different properties for different frequency ranges.

[0145] In many embodiments, each subband is processed by one subband artificial neural network 107 that generates all subband samples for the multi-channel audio signal. However, in some embodiments, one subband includes two or more subband artificial neural networks 107 that generate subband samples for different portions of the subband multi-channel audio signal (as one or more subbands may in fact not include a subband artificial neural network).

[0146] For example, in some embodiments, the downmix audio signal subband samples and the set of upmix parameter values ​​for a given subband are provided to two (or more) subband artificial neural networks 107, which generate subband samples for different parts of the subband of the multi-channel audio signal.

[0147] The two subband artificial neural networks 107 for a given downmix audio signal subband may, for example, generate signals for different time intervals, e.g., one generating subband samples for the first half of the processing time interval and the other generating subband samples for the second half of the processing time interval.

[0148] In other embodiments, one subband artificial neural network 107 generates subband samples for, for example, one subfrequency range, while the other generates samples for another subfrequency range of that subband. Such an approach is particularly suitable for scenarios where, for example, a multi-channel audio signal contains specific tonal components, for example, at specific frequencies. One of the subband artificial neural networks 107 is trained to accurately reflect such tonal components when they are present, while the other subband artificial neural network 107, which does not need to reflect such tonal components, is not impaired by being trained to generate subband samples.

[0149] In such a scenario, each subband artificial neural network is trained specifically to provide output samples for a corresponding portion of the generated subband signal. For example, a subband artificial neural network configured to generate subband samples for the first half of a processing time interval is trained specifically based on a comparison of the generated samples with samples from the first half of the original multi-channel audio signal. Similarly, a subband artificial neural network 107 trained for a particular frequency interval of a subband is trained based on the multi-channel audio signal within that frequency interval.

[0150] The use of such multiple subband artificial neural networks 107 within each (downmix audio signal) subband provides improved performance in many embodiments, often making it possible to generate an improved multi-channel audio signal that more closely corresponds to the original multi-channel audio signal. In particular, it also allows for reduced complexity / computational resources in many embodiments despite the use of more subband artificial neural networks, since subband artificial neural networks can generally each be much less complex and have fewer computational requirements.

[0151] In some embodiments, one or more of the sub-band artificial neural networks 107 are formed by including multiple sub-band artificial neural networks. In particular, as shown in Figure 5, the sub-band artificial neural network 107 is formed by three sub-artificial neural networks. In this case, the sub-band artificial neural network 107 comprises two sub-artificial neural networks 501, 503, both of which receive sub-band samples of the downmix audio signal (corresponding to the sub-band artificial neural network 107 having two input nodes for each sub-band sample). The output nodes of these two sub-artificial neural networks are also input nodes of a third sub-artificial neural network 505, which also has an input node for receiving a set of up-mix parameters. The output node of this third sub-artificial neural network is the output node of the sub-band artificial neural network 107.

[0152] Such an arrangement has been found to be particularly efficient and to provide high-quality reconstruction of multi-channel audio signals. The present approach further allows for sub-training, in particular the first and second sub-artificial neural networks 501, 503, which are individually and separately adapted to provide the desired results. This has been found to be particularly advantageous in many scenarios and for many signals.

[0153] In particular, the first sub-artificial neural network 501 is trained in some embodiments to provide a modified mono signal that is particularly suitable for upmixing, e.g., mono-to-mono processing. Furthermore, the second sub-artificial neural network 503 is trained to provide a decorrelated signal or residual signal for a mono-to-mono downmix audio signal. The third sub-artificial neural network is trained to provide a reconstructed multi-channel audio signal. The third sub-artificial neural network is trained, for example, by end-to-end training based on comparing the original multi-channel audio signal with the reconstructed multi-channel audio signal, based on the first and second sub-artificial neural networks 501, 503 having configurations determined by previous individual trainings.

[0154] In such an approach, the first sub-artificial neural network is generally relatively small because little temporal distortion is generally expected. The second sub-artificial neural network is relatively large, especially at low frequencies, so the artificial neural network better ensures proper decorrelation. The third sub-artificial neural network tends to be relatively small. Overall, reduced complexity and reduced computational resources can generally be achieved.

[0155] Artificial neural networks are adapted to specific purposes through a training process used to adapt / tune / change the weights and other parameters (e.g., biases) of the artificial neural network. It will be appreciated that many different training processes and algorithms for training artificial neural networks are known. Typically, training is based on a large training set, in which a large number of examples of input data are presented to the network. Furthermore, the output of the artificial neural network is typically compared (directly or indirectly) to expected or ideal results. A cost function is generated to reflect the desired outcome of the training process. In a common scenario known as supervised learning, the cost function often represents the distance between the prediction for specific input data and the ground truth. Based on the cost function, the weights are modified, and by repeating the process with the modified weights, the artificial neural network is adapted toward a state where the cost function is minimized.

[0156] More specifically, during the training step, a neural network has two distinct flows of information: from input to output (forward pass) and from output to input (backward pass). In the forward pass, data is processed by the neural network as described above, while in the backward pass, weights are updated to minimize a cost function. Generally, such backward propagation follows the gradient direction of the cost function landscape. In other words, by comparing the predicted output with the ground truth for a batch of data inputs, the direction in which the cost function is minimized and propagated backward can be estimated by updating the weights accordingly. Other known techniques for training artificial neural networks include, for example, the Levenberg-Marquardt algorithm, the conjugate gradient method, and Newton's method.

[0157] In this case, training specifically involves a training set containing a potentially large number of multi-channel audio signals. In some embodiments, the training data are multi-channel audio signals in time segments corresponding to the processing time interval of the artificial neural network being trained, e.g., the number of samples in the training multi-channel audio signals corresponds to the number of samples corresponding to the input nodes of the artificial neural network being trained. Each training example therefore corresponds to one operation of the artificial neural network being trained. Typically, however, batches of training samples are considered for each step to accelerate the training process. Furthermore, many improvements to gradient descent also make it possible to accelerate convergence or avoid local minima in the cost function landscape.

[0158] For each training multi-channel audio signal, the training processor performs a downmix operation to generate a downmix audio signal and corresponding upmix parametric data. Thus, the encoding process applied to the multi-channel audio signals during normal operation is also applied to the training multi-channel audio signals, thereby generating the downmix and upmix parametric data.

[0159] Furthermore, the training processor, in some embodiments, generates a residual signal that reflects differences between the downmix audio signal and the multi-channel audio signal, or more generally, that represents a portion of the multi-channel audio signal that is not adequately represented by the downmix audio signal. For example, in many embodiments, the training processor generates a downmix signal and also generates a residual signal that, when used in upmixing based on the upmix parametric data, results in a (more) accurate reconstructed multi-channel audio signal. Furthermore, the training processor generates upmix parameters.

[0160] In particular, for stereo multi-channel audio signals, the training processor uses a parametric stereo scheme (e.g., according to a suitable standardized method). Such encoding applies frequency- and time-dependent matrix operations, such as rotation operations, to the input stereo signal to generate a downmix signal and a residual signal. For example, a 2×2 matrix / complex-valued multiplication is typically applied to the input stereo signal to substantially align one of the rotated channel signals to have the maximum signal value. This channel is used as a mono signal, and the rotation is typically performed frame-by-frame. The rotation value is stored as part of the upmix parametric data (or a parameter that allows this determination is included in the upmix parametric data). Thus, in the synthesizer, an inverse rotation is performed to reconstruct the stereo signal. Rotating the stereo signal results in another stereo signal, one of whose channels is accordingly aligned with maximum intensity. The other channels are typically discarded in the parametric stereo encoder to reduce the data rate. In conventional parametric stereo decoding, a decorrelated signal is typically generated in the decoder and used for the upmixing process. In current training techniques, this second signal is used as the residual signal for downmixing because it represents information discarded in the encoder, and therefore it represents the ideal signal to be reconstructed in the decoder as part of the upmixing process.

[0161] Thus, in some embodiments, the training processor generates a training downmix signal and / or a training residual signal and / or training upmix parameters from the training multi-channel audio signal.

[0162] The training processor further proceeds to generate subbands for the generated downmix signal (and potentially for the residual signal, if these are used). Similarly, in the case of approaches in which the processing of the parameter generator 105 is not based on an artificial neural network, but is a predetermined operation (e.g., selection or simple interpolation), the training processor further proceeds to generate a set of upmix parameters.

[0163] The training processor thus generates a set of training data including a subband downmix audio signal and a subband set of upmix parameters; in particular, the training processor includes the same encoder and decoder functions (including those of the receiver 101, the subband generator 103, and possibly the parameter generator 105) as those generated by the audio device, resulting in the samples and set of upmix parameters that are fed to the subband artificial neural network 107. Based on these input values, the subband artificial neural network 107 then proceeds to generate subband samples for the multi-channel audio signal, which are transformed into the time domain, resulting in a reconstructed multi-channel audio signal. This reconstructed multi-channel audio signal is then compared to the original training multi-channel audio signal as part of a cost function, which is then used to adapt the subband artificial neural network 107.

[0164] The training is therefore sub-band training, where sub-band data is generated, applied to individual sub-band artificial neural networks 107, and combined with the outputs of the sub-band artificial neural networks 107 into a multi-channel audio signal that can be evaluated by a cost function.

[0165] If the parameter generator 105 also includes one or more artificial neural networks, the generated upmix parametric data is also applied to the parameter generator 105, which generates a set of upmix parameters based on the current configuration. In this case, the training, and in particular the updating, includes the artificial neural networks of the parameter generator 105 as well as the subband artificial neural networks.

[0166] Furthermore, in some embodiments, the cost function includes a contribution that reflects how closely the upmix parameters generated by the artificial neural network of the parameter generator 105 correspond to the original training parameters generated by the training processor. In some embodiments, the parameter generator 105 is trained separately from the subband artificial neural network 107, and the cost function is used based solely on comparing the generated parameter values ​​to the input training upmix parametric data.

[0167] However, in many embodiments, the artificial neural network of the parameter generator 105 is trained together with the subband artificial neural network 107, and the cost function often includes both a contribution indicative of the difference between the input training multi-channel audio signal and the reconstructed multi-channel audio signal, as well as a contribution indicative of the difference between the upmix parameters for the input training multi-channel audio signal and the upmix parameters generated by the artificial neural network of the parameter generator 105.

[0168] 5, the training of the first and / or second sub-artificial neural network 501, 503 is performed separately and before the training of the third sub-artificial neural network 505. For example, the first sub-artificial neural network 501 is trained using the generated training downmix audio signal by comparing it with the resulting downmix audio signal.

[0169] Similarly, the second sub-artificial neural network 503 is trained based on being fed with subband samples of the training downmix audio signal and a cost function based on comparing the resulting output with the generated training residual signal.

[0170] A training approach is used in which outputs from neural network operations are determined from training signals, and a cost function is applied to determine a cost value for each training signal and / or for a combined set of signals (e.g., an average cost value for the training set is determined). The cost function includes various components.

[0171] Generally, the cost function includes at least one component that reflects how close the generated signal is to the reference signal, i.e., the so-called reconstruction error. In some embodiments, the cost function includes at least one component that reflects how close the generated signal is to the reference signal from a perceptual point of view.

[0172] For example, in some embodiments, the multi-channel audio signals generated by the sub-band artificial neural networks (optionally any artificial neural networks of the parameter generator 105) for a given training multi-channel audio signal are compared with the original training multi-channel audio signal. A cost function contribution reflecting the difference between them is generated. This process is performed for all training sets to generate an overall cost function. Furthermore, the technique is applied in each sub-band separately or jointly. In that case, the cost function represents the difference between the sub-bands of the reconstructed multi-channel audio signal and the sub-bands of the original multi-channel audio signal. For example, using independent sub-band artificial neural networks and assuming that perception is not taken into account, training individual sub-band artificial neural networks, for example, using an RMSE-type cost function for each sub-band, is feasible.

[0173] One such example for a stereo multi-channel audio signal is shown in Figure 6. The example in Figure 6 is an example of a training setup for training, in particular, the left channel sub-band artificial neural network 107 and the right channel sub-band artificial neural network 401. (Possibly, one or more artificial neural networks included in the parameter generator 105 are also trained together with the sub-band artificial neural networks using such a training setup.) In scenarios where other artificial neural networks are present, for example, when used in the encoder to generate upmix parametric data (e.g., as a latent representation of the downmix audio signal), it will be understood that such artificial neural networks are added to the shown training setup for joint training.

[0174] In this example, the training processor 601 receives a training multi-channel audio signal, which in this particular example is a stereo signal. In this example, the multi-channel audio signal is received as a sub-band signal, and the following description focuses on the implementation in a single sub-band. The same approach is reused for the other sub-bands.

[0175] For a given training signal, the downmixer 603 performs a downmixing operation to generate a training downmix audio signal, which in the specific example is a training mono audio signal. The downmix audio signal is input to the subband artificial neural network 107.

[0176] Further, the multi-channel audio signal is supplied to a parameter estimator 605, which proceeds to generate up-mix parameters, as is done in an encoder. In this example, the up-mix parameters are generally, for example, IID / IPD / ICC parameters generated by an analysis function applied to the input training multi-channel audio signal, and in particular, the parameters are generated as they are generated in an encoder (e.g., a legacy encoder). The parameters are optionally quantized and coded / decoded in the emulator 607 to generate the up-mix parameters, as is done when they are input to the parameter generator 105.

[0177] The parameter generator 105 and the subband artificial neural network 107, 401 then proceed to reconstruct the stereo multi-channel audio signals l', r'.

[0178] The reconstructed signals l', r' are fed to a comparator 609 which proceeds to generate a cost value based on a cost function that includes a contribution indicative of the difference between the original and reconstructed signals. In many embodiments, the cost function includes a contribution that reflects how closely the upmix parameters generated by the parameter generator 105 match the upmix parameters generated by the parameter estimator.

[0179] It will be appreciated that many different techniques can be used to determine a cost value that reflects the difference between the signals. For example, correlation is performed using a cost value that monotonically decreases as the correlation value increases. As another example, two signals are subtracted from each other and a power measure for the difference signal is used as the cost value. It will be appreciated that many other techniques are available and can be used.

[0180] As a specific example, a training procedure is applied that aims to minimize the distances l-l' and r-r' and to re-instate the upmix parameters (e.g., IID / IPD / ICC) as closely as possible. For a given frame (l, r), adjacent to the energy-preserving downmix m, the (legacy) upmix parameters IID (level difference between the left and right sides per frequency band), ICC (correlation between the left and right sides per frequency band), and IPD (phase difference between the left and right sides) can be determined. Then, the (optional) artificial neural network of the parameter generator 105 and the left and right subband artificial neural networks 107 can be trained together to minimize the reconstruction errors l-l' and r-r' balanced with the recovery of the upmix parameters. This means that the loss function becomes a combination of the signal reconstruction error combined with the PS parameter recovery. In particular, Loss=d reconstruction +α·d stereo and α is a parameter for tuning the balance between signal reconstruction and stereo image reconstruction, where dreconstruction =d(l,l')+d(r,r') Or, in general,

number

number

[0181] Thus, in this example, the cost function generates a cost value that reflects how closely the generated multi-channel audio signal matches the corresponding training multi-channel audio signal.

[0182] Based on the cost values, the training processor 601 adapts the weights of the artificial neural network. For example, a backpropagation technique is used. In particular, the training processor 601 adjusts the weights of both the subband artificial neural network 107 and the artificial neural network of the parameter generator 105 based on the cost values. For example, given the derivative (representing the slope) of the weights with respect to the cost function, the weight values ​​are changed to go in the opposite direction of the gradient. For simple / minimal computation, the case of a single data input backward pass can be referred to as training a perceptron (single neuron).

[0183] This process is repeated until the artificial neural network is considered trained. For example, training is performed for a predetermined number of iterations. As another example, training continues until the weights change less than a predetermined amount. Also very commonly, a validation step is performed in which the network is tested against a validation metric and stopped when it reaches an expected result.

[0184] Artificial neural networks are (further) trained using training data that does not directly represent the audio source / signal, but conveys similar relevant and meaningful information. A particular example is the inclusion of text-based training data. Training artificial neural networks based on text allows the networks to further improve their language understanding and therefore audio reconstruction. For example, by combining text and audio, it becomes easier to predict word sequences than using just one modality. The same applies to video-plus-audio streams (e.g., in the example of lips syncing or lips reading).

[0185] The audio device is in particular implemented in one or more suitably programmed processors. In particular, the artificial neural network is implemented in one or more such suitably programmed processors. The different functional blocks, in particular the artificial neural network, may be implemented in separate processors and / or may, for example, be implemented in the same processor. An example of a suitable processor is provided below.

[0186] 7 is a block diagram illustrating an exemplary processor 700, according to an embodiment of the present disclosure. Processor 700 may be used to implement one or more processors that implement the previously described devices or elements thereof (including, inter alia, one or more artificial neural networks). Processor 700 may be any suitable processor type, including, but not limited to, a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable array (FPGA) programmed to form a processor, an FPGA, a graphic processing unit (GPU), an application specific integrated circuit (ASIC) designed to form a processor, or a combination thereof.

[0187] Processor 700 includes one or more cores 702. Core 702 includes one or more arithmetic logic units (ALUs) 704. In some embodiments, core 702 includes a floating point logic unit (FPLU) 706 and / or a digital signal processing unit (DSPU) 708 in addition to or instead of the ALUs 704.

[0188] The processor 700 includes one or more registers 712 communicatively coupled to the cores 702. The registers 712 are implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 712 are implemented using static memory. The registers provide data, instructions, and addresses to the cores 702.

[0189] In some embodiments, processor 700 includes one or more levels of cache memory 710 communicatively coupled to cores 702. Cache memory 710 provides computer-readable instructions to cores 702 for execution. Cache memory 710 provides data for processing by cores 702. In some embodiments, the computer-readable instructions are provided to cache memory 710 by local memory, for example, local memory attached to external bus 716. Cache memory 710 may be implemented using any suitable cache memory type, for example, metal-oxide-semiconductor (MOS) memory such as static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology.

[0190] Processor 700 includes a controller 714 that controls inputs to processor 700 from other processors and / or components included in the system and / or outputs from processor 700 to other processors and / or components included in the system. Controller 714 controls data paths in ALU 704, FPLU 706, and / or DSPU 708. Controller 714 may be implemented as one or more state machines, data paths, and / or dedicated control logic. Gates in controller 714 may be implemented as standalone gates, FPGAs, ASICs, or any other suitable technology.

[0191] Registers 712 and cache 710 communicate with controller 714 and core 702 via internal connections 720A, 720B, 720C, and 720D. The internal connections may be implemented as buses, multiplexers, crossbar switches, and / or any other suitable connection technology.

[0192] Input and output for processor 700 are provided via bus 716, which may include one or more conductive lines. Bus 716 is communicatively coupled to one or more components of processor 700, such as controller 714, cache 710, and / or registers 712. Bus 716 is coupled to one or more components of the system.

[0193] Bus 716 is coupled to one or more external memories. The external memories include read-only memory (ROM) 732. ROM 732 may be masked ROM, electronically programmable read-only memory (EPROM), or any other suitable technology. The external memories include random access memory (RAM) 733. RAM 733 may be static RAM, battery-backed static RAM, dynamic RAM (DRAM), or any other suitable technology. The external memories include electrically erasable programmable read-only memory (EEPROM®) 735. The external memories include flash memory 734. The external memories include magnetic storage devices, such as disks 736. In some embodiments, the external memories are included within the system.

[0194] The invention may be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention is optionally implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of an embodiment of the invention may be physically, functionally, and logically implemented in any suitable way. Indeed, functionality may be implemented in a single unit, in several units, or as part of other functional units. Thus, the invention may be implemented in a single unit, or may be physically and functionally distributed between different units, circuits, and processors.

[0195] While the present invention has been described with reference to several embodiments, it is not intended that the present invention be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Furthermore, while features may appear to be described with reference to particular embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the terms "comprise" or "include" do not exclude the presence of other elements or steps.

[0196] Furthermore, although individually listed, a plurality of means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Moreover, although individual features may be included in different claims, these may, in some cases, be advantageously combined, and their inclusion in different claims does not imply that the combination of features is infeasible and / or unadvantageous. Furthermore, the inclusion of a feature in one category of claims does not imply limitation to this category, but rather indicates that the feature is equally applicable to other claim categories, where appropriate. Furthermore, the order of features in the claims does not imply that those features must be performed in any particular order, and in particular the order of individual steps in method claims does not imply that those steps must be performed in that order. Rather, steps may be performed in any suitable order. Furthermore, singular references do not exclude pluralities; thus, references to "first," "second," etc. do not exclude pluralities. Reference signs in the claims are provided merely as an illustrative example and shall not be construed in any way as limiting the scope of the claims.

Claims

1. An audio device for generating multi-channel audio signals, A downmixed audio signal for the aforementioned multi-channel audio signal, Upmix parametric data for upmixing the downmix audio signal and A receiver for receiving audio data signals, including, A subband generator for generating a set of frequency subband signals for the subbands of the downmixed audio signal, An artificial neural network configuration comprising multiple subband artificial neural networks, wherein each subband artificial neural network of the multiple subband artificial neural networks generates subband samples for the subbands of the frequency subband representation of the multichannel audio signal, A parameter generator for generating a set of upmix parameter values ​​for subbands of the frequency subband representation of the multichannel audio signal from the upmix parametric data, A generator for generating the multichannel audio signal from the subband samples of the subband of the multichannel audio signal, Equipped with, An audio device in which each subband artificial neural network comprises a set of nodes for receiving a set of upmix parameter values ​​and a sample of at least one frequency subband signal from a set of frequency subband signals, wherein the at least one frequency subband signal is for a subband, and the subband artificial neural network generates subband samples of the multichannel audio signal for that subband.

2. The audio device according to claim 1, wherein at least one first subband artificial neural network of the plurality of subband artificial neural networks comprises a node for receiving parameter values ​​of a set of upmix parameters for subbands other than the subband of the subband artificial neural network.

3. The audio device according to claim 1, wherein at least some parameters of the set of upmix parameter values ​​for different subband artificial neural networks are the same.

4. The audio device according to claim 1, wherein at least some parameters of the set of upmix parameter values ​​differ for different subband artificial neural networks.

5. The audio apparatus according to claim 1, wherein the plurality of subband artificial neural networks comprises separate artificial neural networks for different channels of the multi-channel audio signal for at least one subband.

6. The audio device according to claim 1, wherein the parameter generator modifies the resolution of the set of upmix parameters with respect to the resolution of the upmix parametric data to match the resolution of the processing of the plurality of subband artificial neural networks, the resolution of the processing of the plurality of subband artificial neural networks being one of the frequency resolution of the subbands and the time resolution with respect to the processing time interval for the plurality of subband artificial neural networks.

7. The audio device according to claim 6, comprising at least one artificial neural network, wherein the parameter generator has a node for receiving parameter values ​​of the upmix parametric data and an output node for supplying a set of upmix parameter values ​​for a first subband artificial neural network of the plurality of subband artificial neural networks.

8. The audio apparatus according to claim 1, wherein for at least a first subband, the plurality of subband artificial neural networks comprises at least two subband artificial neural networks that generate samples for different components of the subband signal for the first subband.

9. The audio device according to claim 1, wherein the plurality of subband artificial neural networks are trained with training data having a training input audio signal which includes samples of an input multichannel audio signal, and using a cost function which includes a component that represents the difference between the training input audio signal and the multichannel audio signal generated by the subband artificial neural network.

10. The audio apparatus according to claim 1, wherein the plurality of subband artificial neural networks are trained with training data having a training input audio signal comprising samples of an input multichannel audio signal, and using a cost function that includes a component representing the difference between an upmix parameter for the training input audio signal and an upmix parameter for the multichannel audio signal generated by the subband artificial neural networks.

11. At least one of the multiple subband artificial neural networks is A first sub-artificial neural network having a node that receives samples of frequency subband signals for the subband of the subband artificial neural network, and an output node that supplies samples of a modified downmix audio signal, A second sub-artificial neural network having a node that receives samples of frequency subband signals for the subband of the subband artificial neural network, and an output node that supplies samples of auxiliary audio signals, A third sub-artificial neural network having a node that receives a sample of the modified downmix audio signal, a node that receives a sample of an auxiliary audio signal, and a node that receives a set of upmix parameter values ​​for the subband of the subband artificial neural network, further comprising a third sub-artificial neural network that generates the subband samples for the subband of the frequency subband representation of the multichannel audio signal. The audio device according to claim 1, comprising:

12. The audio apparatus according to claim 1, wherein the set of upmix parameters has a different number of parameters for at least two subbands.

13. The audio device according to claim 1, wherein the upmix parametric data provides parametric data for continuous time intervals, and at least the first subband artificial neural network of the plurality of subband artificial neural networks includes a node for receiving parameter values ​​of a set of upmix parameter values ​​for time intervals of continuous time intervals, which are different from the time intervals in which subband samples of the multichannel audio signal are generated.

14. The audio device according to claim 1, wherein the plurality of subband artificial neural networks do not receive input data other than the subband samples of the downmix audio signal and the parameter values ​​generated from the upmix parametric data.

15. A method for generating a multi-channel audio signal, A downmixed audio signal for the aforementioned multi-channel audio signal, Upmix parametric data for upmixing the downmix audio signal and The steps include receiving an audio data signal, A step of generating a set of frequency subband signals for the subbands of the downmix audio signal, wherein each subband artificial neural network of a plurality of subband artificial neural networks generates subband samples for the subbands of the frequency subband representation of the multichannel audio signal, The steps include generating a set of upmix parameter values ​​for the subbands of the frequency subband representation of the multichannel audio signal from the upmix parametric data, The steps of generating the multichannel audio signal from the subband samples of the subband of the multichannel audio signal, It has, A method comprising a set of nodes, each subband artificial neural network, which receives a set of upmix parameter values ​​and a sample of at least one frequency subband signal from a set of frequency subband signals, wherein the at least one frequency subband signal is for a subband, and the subband artificial neural network generates subband samples of the multichannel audio signal for that subband.

16. A computer program comprising computer program code means for performing all steps of the method according to claim 15 when run on a computer.