Generating multi-channel audio and data signals representing a multi-channel audio signal

JP2025528572A5Pending Publication Date: 2026-08-26KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025514544
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-13
Filing Date
2023-09-06
Publication Date
2026-08-26

AI Technical Summary

Technical Problem

Existing encoding and decoding techniques for multi-channel audio signals, such as parametric stereo, introduce distortions, variations, and artifacts, leading to degraded audio quality and higher data rates, while requiring higher computational resources and complexity, especially on the decoder side.

Method used

Employing a trained artificial neural network to generate an auxiliary audio signal for upmixing, using downmix and upmix parametric data, allowing for improved reconstruction of multi-channel audio signals with reduced complexity and resource usage on the decoder side.

Benefits of technology

The technique provides improved audio quality, reduced complexity, and resource usage on the decoder side, while enabling encoder-side control and adaptation of the decoding process, facilitating efficient generation of multi-channel audio signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The audio device includes a receiver 101 for receiving a data signal including a downmix audio signal for a multi-channel audio signal, upmix parametric data for upmixing the downmix audio signal, and a set of control data values. An artificial neural network 107 has an input node for receiving a second sample of the downmix audio signal and a node for receiving a control data value from the set of control data values. Based on these inputs, the artificial neural network 107 generates a sample of an auxiliary audio signal for the downmix audio signal. A generator 105 generates a multi-channel audio signal from the downmix audio signal and the auxiliary audio signal depending on the upmix parametric data. Another device generates the set of control data values ​​using another artificial neural network having an input node that receives the downmix audio signal of the multi-channel audio signal. In many embodiments, the operation is subband-based, with separate artificial neural networks used for different subbands.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the generation of multi-channel audio signals and / or data signals representing multi-channel audio signals, and particularly, but not exclusively, to the encoding and / or decoding of stereo signals. [Background technology]

[0002] Spatial audio applications are growing in number and popularity, increasingly forming at least part of many audiovisual experiences. Indeed, new and improved spatial experiences and applications are constantly being developed, resulting in increased demands on audio processing and rendering.

[0003] For example, virtual reality (VR) and augmented reality (AR) have been gaining increasing interest in recent years, with numerous implementations and applications reaching the consumer market. Indeed, equipment is being developed for both providing experiences and capturing or recording suitable data for such applications. For example, relatively low-cost equipment is being developed to enable gaming consoles to provide full VR experiences. This trend is expected to continue and, indeed, accelerate as the VR and AR markets reach large scale within a short period of time. In the audio domain, realistic and natural spatial audio reproduction and synthesis are being explored in prominent fields. The ideal goal is to produce a natural sound source such that a user cannot distinguish between the synthesized sound source and the original sound source.

[0004] Many research and development efforts have focused on providing efficient, high-quality audio coding and decoding for spatial audio. Frequently used spatial audio representations are multi-channel audio representations, including stereo representations, and efficient coding of such multi-channel audio has been developed based on downmixing a multi-channel audio signal into downmix channels using fewer channels. One of the major advances in low bitrate audio coding has been the use of parametric multi-channel coding, in which a downmix signal is generated together with parametric data that can be used to upmix the downmix signal to recreate the multi-channel audio signal.

[0005] In particular, instead of traditional mid-side or intensity coding, in parametric multi-channel audio coding, a multi-channel input signal is downmixed to fewer channels (e.g., from two to one) and multi-channel image (stereo) parameters are extracted. The downmix signal is then encoded using a more traditional audio coder (e.g., a mono audio encoder). The downmix bitstream is multiplexed with the coded multi-channel image parameter bitstream. This bitstream is then sent to a decoder, where the process is reversed. First, the downmix audio signal is decoded, and then the multi-channel audio signal guided by the coded multi-channel image / upmix parameters is reconstructed.

[0006] An example of stereo coding is described in E. Schuijers, W. Omen, B. den Brinker, and J. Breebaart, "Advances in Parametric Coding for High-Quality Audio," 114th AES Convention, Amsterdam, The Netherlands, 2003, Preprint 5852. In the described technique, the downmixed mono signal is parameterized by exploiting the natural separation of signals into three components (objects): transient, sinusoidal, and noise. A more detailed description is given in E. Schuijers, J. Breebaart, H. Pumhagen, and J. Engdegard, "Low Complexity Parametric Stereo Coding," 116th AES Convention, Berlin, Germany, 2004, Preprint 6073, which explains how parametric stereo with low (decoder) complexity is achieved when combined with spectral band replication (SBR).

[0007] In the described approach, decoding is based on the use of a so-called decorrelation process, which generates a decorrelated helper signal from the mono signal. In a stereo reconstruction process, both the mono signal and the decorrelated helper signal are used to generate an upmixed stereo signal based on the upmix parameters. In particular, the two signals are multiplied by a time- and frequency-dependent 2x2 matrix with coefficients determined from the upmix parameters to give the output stereo signal.

[0008] However, while parametric stereo (PS) and similar downmix encoding / decoding techniques were a leap from traditional stereo and multi-channel coding, they are not optimal in all scenarios. In particular, known encoding and decoding techniques tend to introduce certain distortions, variations, artifacts, etc. that result in differences between the (original) multi-channel audio signal input to the encoder and the reconstructed multi-channel audio signal at the decoder. Generally, audio quality is degraded and imperfect reconstruction of the multi-channels is performed. Furthermore, data rates are still higher than desired and / or processing complexity / resource usage is higher than preferred.

[0009] A further problem is that it is often particularly desirable to reduce the complexity and computational load on the decoder side, even if this increases the complexity and computational resources on the encoder side. For example, in audio distribution systems, such as systems that provide audiovisual content from a central server to many clients, encoding is performed only once, but decoding is performed many times in individual decoding devices that generally have fewer computational resources than the server device.

[0010] Furthermore, it is often desirable for the encoder side to be able to modify or control the generation of the audio signal at the decoder side, e.g., to be able to adapt the operation of the decoder to provide desired audio characteristics, or to be able to adapt the processing to provide a more accurate reconstruction of the multi-channel audio signal at the decoder side. Summary of the Invention [Problem to be solved by the invention]

[0011] Therefore, improved techniques would be advantageous, particularly techniques that allow for increased flexibility, improved adaptability, improved performance, improved audio quality, trading off improved audio quality against data rate, reduced complexity and / or resource usage, improved remote control of audio processing, improved encoder-side input to decoder-side operation / processing, improved computational resource allocation between encoder-side and decoder-side processing, reduced computational burden, facilitated implementation, and / or an improved spatial audio experience.

[0012] SUMMARY OF THE INVENTION Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination. [Means for solving the problem]

[0013] According to one aspect of the present invention, there is provided an apparatus for generating a multi-channel audio signal, the apparatus comprising: a receiver (101) for receiving a data signal comprising a downmix audio signal for the multi-channel audio signal, upmix parametric data for upmixing the downmix audio signal, and a set of control data values ​​for adapting a process for generating an auxiliary signal for upmixing from the downmix audio signal; an artificial neural network (107) having input nodes for receiving samples of the downmix audio signal and output nodes for supplying samples of the auxiliary audio signal, the artificial neural network (107) further comprising nodes arranged for including a contribution from the set of control data values; and a generator (105) for generating the multi-channel audio signal from the downmix audio signal and the auxiliary signal in dependence on the upmix parametric data.

[0014] The present technique provides an improved audio experience in many embodiments, and for many signals and scenarios, the technique provides improved generation / reconstruction of multi-channel audio signals, as well as improved perceived audio quality.

[0015] The present approach provides a particularly advantageous arrangement that in many embodiments and scenarios allows for facilitating and / or improving the possibilities of utilizing artificial neural networks in audio processing, including audio encoding and / or decoding in general, and allows for the advantageous employment of artificial neural networks in generating a multi-channel audio signal from a downmix audio signal.

[0016] The present technique, in many embodiments, enables an encoder to adapt / modify the processing of a device that generates the multi-channel audio signal, thereby enabling improved multi-channel audio signals to be generated. The present technique enables or facilitates such remote control by providing specific techniques that accommodate and exploit the advantages that may be achieved by employing artificial neural networks as part of the audio processing.

[0017] The technique provides an efficient implementation and in many embodiments allows for reduced complexity and / or resource usage. The technique allows for a reduction in data rate for data representing a multi-channel audio signal using a downmix signal in many scenarios. The technique in many embodiments allows for reduced complexity and / or resource usage at the decoder / generation / reconstruction side, at the expense of potentially increased complexity and / or resource usage at the encoder / source side.

[0018] The samples of the downmix audio signal are time-domain samples or frequency-domain samples (particularly sub-band samples). The samples of the auxiliary audio signal are time-domain samples or frequency-domain samples (particularly sub-band samples). The samples may span a particular time and frequency range.

[0019] The upmix parametric data includes parameters (values) relating properties of the downmix signal to properties of the multi-channel audio signal. The upmix parametric data includes data indicative of relative properties between channels of a multi-channel audio signal. The upmix parametric data includes data indicative of differences in properties between channels of a multi-channel audio signal. The upmix parametric data includes data that is perceptually relevant to the synthesis of a multi-channel audio signal. Properties are, for example, differences in phase and / or magnitude and / or timing and / or correlation. In some embodiments and scenarios, the upmix parametric data expresses abstract properties that cannot be directly understood by humans / experts (but generally facilitate better reconstruction / lower data rates, etc.). The upmix parametric data includes data including at least one of inter-channel magnitude difference, inter-channel timing difference, inter-channel correlation and / or inter-channel phase difference for the channels of a multi-channel audio signal.

[0020] An artificial neural network is a trained artificial neural network.

[0021] The artificial neural network is a trained artificial neural network trained with training data including a training downmix audio signal and training upmix parametric data generated from the training multi-channel audio signal, where the training employs a cost function that compares the training multi-channel audio signal to an upmixed multi-channel signal generated from the training downmix audio signal and the generated auxiliary audio signal using the training upmix parametric data. The artificial neural network is a trained artificial neural network trained with training data including training data representing a range of relevant audio sources, including video, motion picture, telecommunication, etc. recordings.

[0022] The artificial neural network is an artificial neural network trained with training data having training input data including a training downmix audio signal of the training multi-channel audio signal, and using a cost function including a contribution indicative of a difference between a training auxiliary audio signal generated by a second artificial neural network in response to the training data and a training residual signal for the training downmix audio signal.

[0023] The generator is configured to generate the multi-channel audio signal by applying a matrix multiplication to the downmix signal and the auxiliary audio signal with coefficients of a matrix determined as a function of parameters of the upmix parametric data, the matrix being time and frequency dependent.

[0024] The audio device is in particular an audio decoder device.

[0025] The control data values ​​may be referred to as metadata, conditioning features, latent expressions, conditioning variables, and / or composite numbers.

[0026] According to an optional feature of the invention, the apparatus comprises a filter bank (401) for generating a frequency subband representation of a downmix audio signal, the samples of the downmix audio signal being subband samples of the frequency subband representation.

[0027] Sub-band processing provides particularly advantageous performance in many embodiments, and the present arrangement is particularly well suited to sub-band processing, which allows for reduced complexity and / or improved multi-channel audio signals to be generated.

[0028] According to an optional feature of the invention, the artificial neural network (107) is an artificial neural network of a first plurality of subband artificial neural networks, each subband artificial neural network of the first plurality of subband neural networks generating subband samples for a subset of subbands of the frequency subband representation of the auxiliary audio signal.

[0029] A particular advantage of this approach is that it allows for efficient sub-band processing, thereby enabling the required processing to be split across multiple smaller artificial neural networks, which generally allows for reduced complexity and / or improved multi-channel audio signals to be produced.

[0030] In some embodiments, the first plurality of subband artificial neural networks includes an artificial neural network for each subband of the frequency subband representation of the auxiliary audio signal.

[0031] In some embodiments, the number of input nodes for the neural networks of the first plurality of sub-band neural networks decreases monotonically with increasing frequency.

[0032] In some embodiments, the generator is configured to generate the frequency / sub-band representation of the multi-channel signal by performing upmixing on the frequency / sub-band representation of the auxiliary audio signal and on the frequency / sub-band representation of the downmix audio signal, and to convert the multi-channel signal frequency / sub-band representation of the multi-channel signal into a time-domain representation of the multi-channel signal.

[0033] In many embodiments, each (or at least some) subband neural network is configured to generate subband samples for one subband of the frequency subband representation of the auxiliary audio signal.

[0034] According to an optional feature of the invention, at least one control data value of the set of control data values ​​is processed by at least two artificial neural networks of the first plurality of subband artificial neural networks.

[0035] This provides for very advantageous and efficient implementation and / or operation and / or performance in many embodiments and scenarios.

[0036] In some embodiments, at least some of the sets of control data values ​​are common for multiple subbands of the frequency subband representation of the downmix audio signal.

[0037] According to an optional feature of the invention, at least one control data value of the set of control data values ​​is not processed by at least one artificial neural network of the first plurality of subband artificial neural networks.

[0038] This provides for very advantageous and efficient implementation and / or operation and / or performance in many embodiments and scenarios.

[0039] According to an optional feature of the invention, the artificial neural network (107) is trained with training data including at least one of a training multi-channel audio signal, a training downmix audio signal, and training upmix parametric data generated from the training multi-channel audio signal, wherein the training employs a cost function that compares the training multi-channel audio signal with an upmixed multi-channel signal generated from the training downmix signal and from the generated auxiliary audio signal using the training upmix parametric data.

[0040] This provides for particularly efficient implementation and / or improved performance, which in many embodiments provides for particularly efficient and high performance training.

[0041] The set of control data values ​​generated by the sub-band neural networks and common to the plurality of sub-bands is input to a plurality of sub-band artificial neural networks that generate the multi-channel audio signal (in particular, the second artificial neural network is one of the plurality of artificial neural networks having nodes that receive the common set of control data values).

[0042] The artificial neural network is trained with training data including the downmix signal and upmix parametric data generated from the multi-channel audio signal and control data values, and the training employs a cost function that compares the training multi-channel audio signal with an upmixed multi-channel signal resulting from upmixing the training downmix signal and an auxiliary signal at an output by the artificial neural network using the training upmix parametric data.

[0043] According to an optional feature of the invention, the artificial neural network (107) is trained with training data having training input data including a training downmix audio signal of the training multi-channel audio signal and a training set of control data values ​​for the training downmix audio signal, and using a cost function including a contribution indicative of a difference between a training auxiliary audio signal generated by the artificial neural network in response to the training data and a training residual signal for the training downmix audio signal.

[0044] This provides for particularly efficient implementation and / or improved performance, which in many embodiments provides for particularly efficient and high performance training.

[0045] According to an optional feature of the invention, the cost function further comprises a contribution indicative of the degree of correlation between the auxiliary audio signal and a downmix audio signal of the multi-channel audio signal.

[0046] This provides for particularly efficient implementation and / or improved performance, which in many embodiments provides for particularly efficient and high performance training.

[0047] According to an optional feature of the invention, the apparatus further comprises an upsampler for temporally upsampling the set of control data values, the artificial neural network (107) including contributions from the upsampled control data values.

[0048] This provides a particularly efficient implementation and / or improved performance.

[0049] According to one aspect of the present invention, there is provided an apparatus for generating a data signal representing a multi-channel audio signal, the apparatus comprising: a downmixer for generating a downmix audio signal by downmixing the multi-channel audio signal, the downmixer further generating upmix parametric data for upmixing the downmix audio signal; an artificial neural network for generating a set of control data values ​​for adapting a process for generating an auxiliary audio signal for upmixing from the downmix audio signal, the artificial neural network having output nodes for supplying the set of control data values ​​and input nodes for receiving samples of at least one of the downmix audio signal and the multi-channel audio signal; and a generator for generating a data signal to include the downmix audio signal, the upmix parametric data, and the set of control data values.

[0050] The technique provides, in many embodiments, an improved audio experience. For many signals and scenarios, the technique provides improved generation / reconstruction of multi-channel audio signals, along with improved perceived audio quality. The technique provides, in many embodiments and scenarios, a particularly advantageous arrangement that allows for facilitating and / or improving the possibility of utilizing artificial neural networks in audio processing, including audio encoding and / or decoding in general. The technique allows for the advantageous employment of artificial neural networks in generating multi-channel audio signals from downmix audio signals.

[0051] The technique provides an efficient implementation and allows for reduced complexity and / or resource usage in many embodiments. The technique allows for a reduction in data rate for data representing a multi-channel audio signal using a downmix signal in many scenarios. The technique allows for reduced complexity and / or resource usage at the decoder / generation / reconstruction side in many embodiments, at the expense of potentially increased complexity and / or resource usage at the encoder / source side.

[0052] The remarks given above in relation to an apparatus for generating a multi-channel audio signal apply mutatis mutandis to an apparatus for generating a data signal.

[0053] The samples of the downmix audio signal may be time domain samples or frequency domain samples (particularly subband domain samples). The samples of the auxiliary audio signal may be time domain samples or frequency domain samples (particularly subband domain samples).

[0054] An artificial neural network is a trained artificial neural network.

[0055] The artificial neural network is a trained artificial neural network trained with training data including a training downmix audio signal and training upmix parametric data generated from the training multi-channel audio signal, wherein the training employs a cost function that compares the training multi-channel audio signal with an upmixed multi-channel signal generated from the training downmix audio signal and the generated auxiliary audio signal using the training upmix parametric data.

[0056] The artificial neural network is a trained artificial neural network trained with training data having training input data including a training downmix audio signal of the training multi-channel audio signal and using a cost function including a contribution indicative of a difference between a training auxiliary audio signal generated by a second artificial neural network in response to the training data and a training residual signal for the training downmix audio signal.

[0057] The audio device is in particular an audio encoder device.

[0058] The control data values ​​may be referred to as metadata, conditioning features, latent expressions, conditioning variables, and / or composite numbers.

[0059] According to an optional feature of the invention, the set of control data values ​​provides a latent representation of at least one of a downmix audio signal and a multi-channel audio signal.

[0060] This provides for a particularly efficient implementation and / or improved performance.

[0061] According to an optional feature of the invention, the artificial neural network is trained with training data having training input data that are samples of a training audio signal, the training audio signal being at least one of a multi-channel audio signal and a downmix audio signal of the multi-channel audio signal, and using a cost function that includes a component indicative of a difference between the training audio signal and an output audio signal generated from a set of control data values ​​generated by the artificial neural network for the training audio signal.

[0062] This provides for particularly efficient implementation and / or improved performance, which in many embodiments provides for particularly efficient and high performance training.

[0063] According to an optional feature of the invention, the output audio signal is the output of a composite artificial neural network having input nodes that receive a set of control data values ​​generated by the artificial neural network for a training audio signal, the artificial neural network and the composite artificial neural network being trained together.

[0064] This provides for particularly efficient implementation and / or improved performance, which in many embodiments provides for particularly efficient and high performance training.

[0065] According to an optional feature of the invention, the training audio signal and the output audio signal are downmix audio channels.

[0066] According to one aspect of the present invention, there is provided a method for generating a multi-channel audio signal, the method comprising the steps of receiving a data signal comprising a downmix audio signal for the multi-channel signal, upmix parametric data for upmixing the downmix audio signal, and a set of control data values ​​for adapting a process for generating an auxiliary signal for upmixing from the downmix audio signal; an artificial neural network (107) having input nodes for receiving samples of the downmix audio signal and output nodes for supplying samples of the auxiliary audio signal, the artificial neural network (107) further comprising nodes arranged for including contributions from the set of control data values; and generating the multi-channel audio signal from the downmix audio signal and the auxiliary signal in dependence on the upmix parametric data.

[0067] According to one aspect of the present invention, there is provided a method for generating a data signal representing a multi-channel audio signal, the method comprising the steps of: generating a downmix audio signal by downmixing the multi-channel audio signal, the downmixer further generating upmix parametric data for upmixing the downmix audio signal; generating a downmix audio signal; an artificial neural network generating a set of control data values ​​for adapting processing for generating an auxiliary audio signal for upmixing from the downmix audio signal, the artificial neural network having output nodes for supplying the set of control data values ​​and input nodes for receiving samples of at least one of the downmix audio signal and the multi-channel audio signal; and generating a data signal to include the downmix audio signal, the upmix parametric data, and the set of control data values.

[0068] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter.

[0069] Embodiments of the present invention will now be described, by way of example only, with reference to the drawings in which: [Brief explanation of the drawings]

[0070] [Figure 1] FIG. 2 illustrates some elements of an example audio device, according to some embodiments of the present invention. [Figure 2] FIG. 1 illustrates an example of the structure of an artificial neural network. [Figure 3] FIG. 1 illustrates an example of a node of an artificial neural network. [Figure 4] FIG. 2 illustrates some elements of an example audio device, according to some embodiments of the present invention. [Figure 5]FIG. 2 illustrates some elements of an example audio device, according to some embodiments of the present invention. [Figure 6] FIG. 2 illustrates some elements of an example audio device, according to some embodiments of the present invention. [Figure 7] FIG. 1 illustrates some elements of an example apparatus for training an artificial neural network for an audio device, according to some embodiments of the present invention. [Figure 8] 2A-2C illustrate some elements of possible configurations of a processor for implementing elements of an audio device according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0071] FIG. 1 illustrates some elements of an audio device according to some embodiments of the present invention.

[0072] The audio device comprises a receiver 101 configured to receive a data signal / bitstream comprising a downmix audio signal that is a downmix of a multi-channel audio signal. The following description focuses on the case where the multi-channel audio signal is a stereo signal and the downmix signal is a mono signal, but it will be understood that the techniques and principles described are equally applicable to multi-channel audio signals having more than two channels, and to downmix signals having more than two channels (although having fewer channels than a multi-channel audio signal).

[0073] Furthermore, the received data signal includes upmix parametric data for upmixing the downmix audio signal. The upmix parametric data is a set of parameters that indicate, in particular, a relationship between signals of different audio channels of a multi-channel audio signal (in particular a stereo signal) and / or between the downmix signal and an audio channel of the multi-channel audio signal. Typically, the upmix parameters indicate a measure of similarity, such as a time difference, a phase difference, a level / intensity difference, and / or a correlation. Typically, the upmix parameters are provided per time and per frequency (time-frequency tile). For example, new parameters for a set of subbands are provided periodically. The parameters include, in particular, an inter-channel phase difference (IPD) parameter, an overall phase difference (OPD) parameter, an inter-channel correlation (ICC) parameter, and a channel phase difference (CPD) parameter, which are known from parametric stereo coding (as well as from higher channel coding).

[0074] Typically, the downmix audio signal is coded and the receiver 101 includes a decoder 103 for decoding the downmix audio signal, i.e., in the particular example, a mono signal. It will be understood that the decoder 103 is not required if the received downmix audio signal is not coded and that the decoder 103 is considered to be an integral part of the receiver 101.

[0075] The receiver 101 is coupled to a generator 105, which generates a multi-channel audio signal from the downmix signal. The generator 105 is configured to generate a multi-channel audio signal from the downmix audio signal and from the auxiliary audio signal depending on the parametric upmix data. The generator, in particular for the stereo case, generates an output multi-channel audio signal by applying a time- and frequency-dependent 2×2 matrix multiplication to samples of the downmix audio signal and the auxiliary audio signal. The coefficients of the 2×2 matrix are generally determined on a time- and frequency-band basis from upmix parameters of the upmix parametric data. For other upmix operations, such as from a mono or stereo downmix signal to a five-channel multi-channel audio signal, the generator 105 applies a matrix multiplication using a matrix of suitable dimensions.

[0076] It will be appreciated that many different techniques are known to those skilled in the art for generating such a multi-channel audio signal from a downmix audio signal and an auxiliary audio signal, and for determining suitable matrix coefficients from upmix parametric data, and that any suitable technique may be used. In particular, various techniques for parametric stereo upmixing based on a downmix and an auxiliary audio signal are well known to those skilled in the art.

[0077] In conventional systems, upmixing involves generating an auxiliary audio signal in the form of a decorrelated signal of the mono audio signal. It has been found that generating a decorrelated signal and mixing it with the mono audio signal results in a perceived improvement in the quality of the upmix signal, and decoders have therefore been developed to take advantage of this. The decorrelated signal is typically generated by a decorrelator in the form of an all-phase filter applied to the mono audio signal. However, while the use of such an all-pass filter tends to result in a multi-channel audio signal of perceived improved quality, it is still not ideal, and some degradation in audio quality is often perceived.

[0078] The audio device of Figure 1 uses techniques that have been found to tend to provide improved perceived audio quality in many scenarios and for many different audio signals.

[0079] In this approach, the decorrelated signal is not generated by simple filtering of the downmix / mono audio signal, but rather an auxiliary audio signal is generated by a trained artificial neural network, which is used by the generator 105 to generate a multi-channel audio signal based on the upmix parameters.

[0080] The device further comprises an artificial neural network, also referred to as synthesis artificial neural network 107. The synthesis artificial neural network 107 is coupled to the decoder 103 / receiver 105. The synthesis artificial neural network 107 comprises, in particular, an input node for receiving samples of the downmix audio signal. The synthesis artificial neural network 107 has an output node for providing samples of the auxiliary audio signal, which are then fed to the generator 105, where the upmix operation is completed.

[0081] In this approach, the upmix process is therefore not based on applying a decorrelation filter to the downmix audio signal to generate a decorrelated signal that is then combined with the downmix audio signal to generate a multi-channel audio signal. Rather, a trained artificial neural network generates an auxiliary audio signal that specifically replaces the decorrelated signal used in conventional upmix decoders.

[0082] As will be explained in more detail later, different approaches are used to train the artificial neural network. In particular, holistic training seeks to result in a multi-channel audio signal output from an audio device that most closely corresponds to the original multi-channel audio signal. Thus, the synthesis artificial neural network 107 is trained to provide an auxiliary audio signal that most effectively results in an accurate reconstruction of the multi-channel audio signal. In contrast to conventional approaches, such a signal is not necessarily a decorrelated signal; rather, the synthesis artificial neural network 107 is trained to provide an auxiliary audio signal that is most suitable for combination with a downmix audio signal to generate a multi-channel audio signal. Such a signal is generally not a decorrelated signal of the downmix audio signal, but, for example, a partially correlated signal, which may in fact often be closer to the actual residual signal resulting from the original downmix of the multi-channel audio signal. Thus, the user of the trained artificial neural network allows the decoder to implicitly and automatically take into account and compensate for effects introduced at the encoder side.

[0083] In this approach, the artificial neural network configuration is therefore configured to generate a second, auxiliary "helper" signal that assists and improves the multi-channel reconstruction. For example, in the case of a stereo signal, the encoder generates c *The downmix signal is generated as (l+r), where l and r represent the left and right channel signals, respectively, and c represents a time- and frequency-dependent scaling factor. The corresponding second signal for ideal reconstruction is d * (lr), where d is again time and frequency dependent. The two signals are not necessarily perfectly decorrelated, and a substantial advantage of the described approach is that, as opposed to simply attempting to decorrelate a mono downmix, the artificial neural network architecture can map the ideal signal d * The goal is to generate auxiliary audio signals that tend to approximate (lr), which generally results in a significantly improved reconstruction of the original multi-channel audio signal.

[0084] The generation of the auxiliary audio signal is further improved by a synthesis artificial neural network 107, which generates a signal adapted based on a set of control data values ​​received together with the downmix audio signal. In this approach, the synthesis artificial neural network 107 further comprises nodes that receive contributions from control values ​​of the set of control data values ​​received by the receiver 101, typically provided to the synthesis device from a source device. Such nodes may be input nodes of an input layer that also comprises nodes that receive samples of the downmix audio signal, or one or more or all of the nodes receiving the control data values ​​may be nodes of a layer different from the input layer of the downmix audio signal. For example, some or all of the nodes receiving contributions from the control data values ​​are part of a hidden layer or a processing layer of the synthesis artificial neural network 107.

[0085] In this approach, specific configurations are accordingly used to enable the processing and operation of the synthesis device to be remotely modified / adapted, in particular a technique that enables encoder-side control of decoder-side reconstruction of a multi-channel audio signal. The approach provides, in particular, a way for encoder-side control of the generation of auxiliary audio signals and for adapting the operation of a trained local artificial neural network that performs an integral part of the synthesis of a multi-channel audio signal. The approach enables, for example, an encoder device to generate signal-specific / dependent control data that can adapt the operation of the decoder to generate an improved auxiliary audio signal. Furthermore, the approach enables remote control of functions that are not known or unpredictable at the encoder / source side, or indeed even at the decoder / synthesis side. The approach enables, in particular, encoder / remote control of an artificial neural network so that such encoder / remote control can be adapted to provide more desirable or improved operation. Furthermore, this approach allows, in some embodiments, the complexity of the synthesizer to be reduced, as signal-specific control data can be generated at the source, rather than requiring the synthesizer to process the signal to generate data that can adapt the second artificial neural network 109, for example.

[0086] The present approach, in many embodiments, allows for improved reconstruction of multi-channel audio signals and provides a highly efficient implementation achieved with relatively low complexity and resource usage.

[0087] The artificial neural network used in the described functions is a network of nodes organized into layers, each node having a node value. Figure 2 shows an example of a section of an artificial neural network.

[0088] The node value for a given node is calculated to include contributions from some, or often all, nodes in previous layers of the artificial neural network. In particular, the node value for a node is calculated as a weighted sum of the node values ​​of all node outputs in the previous layer. Typically, a bias is added, and the result is applied to an activation function. The activation function typically accounts for the essential part of each neuron by providing nonlinearity. Such nonlinearity and activation function have a significant effect on the learning and adaptation process of the neural network. Therefore, the node value is generated as a function of the node values ​​in the previous layer.

[0089] The artificial neural network comprises, inter alia, an input layer 201 comprising a plurality of nodes that receive input data values ​​for the artificial neural network. Thus, the node values ​​for the nodes of the input layer generally become input data values ​​to the artificial neural network directly and are therefore not calculated from other node values.

[0090] The artificial neural network may further comprise zero, one or more hidden layers 203 or processing layers. For each such layer, the node values ​​are generally generated depending on the node values ​​of the nodes in the previous layer, in particular depending on a weighted combination and an additional bias followed by an activation function (such as a sigmoid, ReLU or Tanh function applied).

[0091] In particular, as shown in Figure 3, each node, sometimes called a neuron, receives input values ​​(from nodes in previous layers) and then calculates the node value as a function of these values. Often, this involves first generating the value as a linear combination of the input values, where each input value is represented by a weight

number

[0092] An activation function is then applied to the resulting combination. For example, a node value of 1 is l=f(k) where the function is the Rectified Linear Unit function, as described, for example, in Xavier Glorot, Antoine Bordes, Yoshua Bengio, Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, PMLR 15:315-323, 2011), i.e., f(k) = ReLU(k) = max(0,k) is.

[0093] Other frequently used functions include the sigmoid function or the tanh function. In many embodiments, node outputs or values ​​are calculated using multiple functions. For example, f(k)=ReLU(k)=δ(k) Both ReLU and sigmoid functions are combined using activation functions such as

[0094] Such operations are performed by each node of the artificial neural network (generally except for the input nodes).

[0095] The artificial neural network further comprises an output layer 205 that provides the output from the artificial neural network, i.e., the output data of the artificial neural network are the node values ​​of the output layer. For hidden / processing layers, the output node values ​​are generated by functions of the node values ​​of the previous layer. However, in contrast to the hidden / processing layers, where the node values ​​are generally not accessible or further used, the node values ​​of the output layer are accessible and provide the results of the operation of the artificial neural network.

[0096] Several different network structures and toolboxes for artificial neural networks have been developed, and in many embodiments, artificial neural networks are based on adapting and customizing such networks. One example of a network architecture suitable for the above-mentioned applications is WaveNet by van den Oord et al., described in Oord, Aaron van den, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. "Wavenet: A generative model for raw audio." arXiv preprint arXiv:1609.03499 (2016).

[0097] WaveNet is an architecture used for synthesis of time-domain signals using dilated causal convolution, and has been successfully applied to audio signals. In WaveNet, the following activation function is used:

number

number

[0098] Artificial neural networks are sometimes further configured to include additional contributions that allow the artificial neural network to be dynamically adapted or customized for particular desired properties or characteristics of the generated output. For example, a set of values ​​is provided to adapt the artificial neural network. These values ​​are included by providing contributions to some nodes of the artificial neural network. These nodes are specifically input nodes, but generally nodes in hidden or processing layers. Such adaptation values ​​are weighted and added, for example, as contributions to a weighted sum / correlation value for a given node. For example, in the case of WaveNet, such adaptation values ​​are included in an activation function. For example, the output of the activation function is

number

[0099] The above description relates to a neural network approach that is suitable for many embodiments and implementations. However, it will be understood that many other types and structures of neural networks may be used. Indeed, many different approaches for generating neural networks have been and are being developed, including neural networks that use complex structures and processes different from those described above. The approach is not limited to any particular neural network approach, and any suitable approach may be used without detracting from the invention.

[0100] In many embodiments, the set of control data values ​​is generated by an audio device, also referred to as an encoder or source audio device. The source audio device provides a bitstream including a data signal, in particular a downmix audio signal, upmix parametric data, and a set of control data values. The source audio device is in particular an encoder device that receives a multi-channel audio signal and encodes the multi-channel audio signal by generating a downmix audio signal with associated upmix parametric data. In particular, the source audio device is a parametric stereo (PS) encoder that receives a stereo signal and encodes it as a mono audio signal with associated upmix parametric data.

[0101] Furthermore, the source audio device is configured to include in the bitstream a set of control data values ​​to enable the synthesis audio device to adapt its upmixing process. In particular, the source audio device generates and transmits a set of control data values ​​that enable it to adapt the upmixing process of the synthesis audio device, and in particular the source audio device may adapt the operation of an artificial neural network that generates the auxiliary audio signal in the synthesis audio device.

[0102] Different approaches may be used by different source audio devices to generate the set of control data values. Indeed, in some embodiments, the source audio device stores a predetermined set of control data values ​​that adapt the upmixing operation, and in particular the operation of the artificial neural network of the synthesis audio device. This allows, for example, different source audio devices to adapt the operation of the synthesis audio device to their individual preferences. As another example, this approach allows for adaptation of the audio system without requiring the updating of all end-user devices. For example, if improved training or adaptation to different types of signals is to be brought into the system, this is achieved by changing the set of control data values ​​stored in the source audio device. This allows the synthesis artificial neural network to be updated to perform improved operations without requiring retraining or modification of the synthesis artificial neural network itself.

[0103] In many embodiments, the source audio device is configured to determine the set of control data values ​​in dependence on the multi-channel audio signal or the generated downmix audio signal.

[0104] In some embodiments, one or more of the control data values ​​are generated by analysis of the downmix audio signal.

[0105] It will be appreciated that many different algorithms and procedures are known for analysing audio signals to extract or determine properties of the audio signals, and that any suitable (analysis) techniques, algorithms, and properties may be used.

[0106] In many embodiments, the source audio device includes functionality for generating the set of control data values ​​based on an artificial neural network that generates the set of control data values ​​from the downmix audio signal. An example of such an encoder is shown in Figure 4.

[0107] The source audio device of Fig. 4 comprises a downmixer 401 which receives a multi-channel audio signal, in this case a stereo signal. The downmixer 401 generates a downmix audio signal, in a particular example a mono signal. Furthermore, the downmixer 401 generates upmix parameters, such as IPD parameters, OPD parameters, ICC parameters, CPD parameters, etc. The downmixer 401 further encodes the downmix audio signal in a suitable format. It will be understood that many different algorithms and techniques are known for generating a downmix audio signal with associated upmix parameters, and that any suitable technique may be used (e.g., PS coding is applied in this case).

[0108] The (encoded) downmix audio signal and the upmix parameters are fed to a data signal generator 403 configured to generate an output data signal / bitstream comprising the (encoded) downmix audio signal and upmix parametric data comprising the upmix parameters.

[0109] The source audio device further comprises an artificial neural network, hereafter referred to as source artificial neural network 405. The source artificial neural network 405 receives samples of the downmix audio signal and / or the multi-channel audio signal and in response generates a set of control data values. Such set of control data values ​​is supplied to the data signal generator 403, which includes them in the generated bitstream.

[0110] The source audio device therefore comprises, in particular, a source artificial neural network 405 configured to receive samples of the downmix audio signal and / or the multi-channel audio signal, which samples are fed to input nodes of the source artificial neural network 405. Output nodes of the source artificial neural network 405 provide a set of control data values ​​for the downmix audio signal / multi-channel audio signal. The set of control data values ​​comprises several values ​​(in particular scalar values) reflecting properties or characteristics of the downmix audio signal / multi-channel audio signal.

[0111] Typically, the source artificial neural network 405 has a much larger number of input nodes than output nodes, so that a relatively small number of control data values ​​are generated from a relatively large number of samples. The control data values ​​thus provide a highly compressed and reduced representation of some properties or characteristics of the downmix audio signal / multi-channel audio signal. For example, the source artificial neural network 405 has 1028 input nodes and 16 output nodes, which provides a highly compressed set of values ​​that depend on the properties of the downmix audio signal. A latent representation of the downmix audio signal and / or multi-channel audio signal is generated.

[0112] In this configuration, the output of the source artificial neural networks 405 is provided to the synthesized artificial neural network 107 (via data signal distribution / communication) and is therefore used to control and adapt the processing of the synthesized artificial neural network 107. Thus, even with constant weights and coefficients, the synthesized artificial neural network 107 can be viewed as not simply a fixed filter or trained network, but rather as an adaptive or variable operation that is adapted based on the results of the source artificial neural networks 405.

[0113] Similarly, the source artificial neural network 405 and the generated control data values ​​are not specifically control data values ​​that represent a particular selected property or characteristic of the signal. Rather, the source artificial neural network 405 is trained so that it automatically adapts to provide control data values ​​that are particularly suitable for adapting the synthesized artificial neural network 107 to provide improved output values ​​for accurate reproduction of the multi-channel audio signal.

[0114] Such an approach generally provides a particularly efficient approach leading to an improved reconstruction of the multi-channel audio signal at the synthesis side. The approach therefore generates a downmix audio signal and / or a latent representation of the multi-channel audio signal.

[0115] In many embodiments, the synthesis audio device is configured to perform subband processing. In particular, as shown in Figure 5, the device of Figure 1 is modified to include a filter bank configured to generate frequency subband representations of the downmix audio signal. The filter bank may be a quadrature mirror filter (QMF) bank and may be implemented, for example, by a fast Fourier transform (FFT), although it will be understood that many other filter banks and techniques for splitting an audio signal into multiple subband signals are known and may be used. The filter bank is in particular a complex-valued pseudo-QMF bank, resulting in, for example, 32 or 64 complex-valued subband signals.

[0116] In many embodiments, the filter bank 501 is configured to generate a set of subband signals for subbands having equal bandwidths. In other embodiments, the filter bank 401 is configured to generate subband signals using subbands having different bandwidths. For example, higher frequency subbands have higher bandwidths than lower frequency subbands. Also, subbands are grouped together to form higher bandwidth subbands.

[0117] Generally, the subbands have bandwidths in the range of 10 Hz to 10,000 Hz.

[0118] In some such embodiments, the artificial neural network that generates the samples for the auxiliary audio signal is configured to receive subband samples, i.e., samples of the subband audio signal. In particular, the apparatus of Figure 5 comprises a plurality of subband neural networks 109, each configured to receive subband samples for a subband generated by the filter bank from the downmix signal. Each of these subband artificial neural networks further receives a set of control data values, which proceed to generate subband samples of the multi-channel audio signal for that subband.

[0119] The subband samples from the subband artificial neural network are provided to generator 105, which proceeds to generate a reconstructed multi-channel audio signal. For example, in some embodiments where a subband representation of the multi-channel audio signal is desired (e.g., with subsequent processing that is also subband-based), generator 105 simply outputs the subband samples from the subband artificial neural network, possibly according to a particular structure or format. In many embodiments, generator 105 comprises functionality for converting the subband representation of the reconstructed multi-channel audio signal into a time-domain representation. Generator 105 particularly comprises a synthesis filterbank that performs the inverse operation of filterbank 501, thereby converting the subband representation into a time-domain representation of the multi-channel audio signal.

[0120] The generator is particularly configured to generate the frequency / subband domain representation of the multi-channel audio signal by processing the frequency or subband domain representation of the downmix audio signal and the frequency / subband domain representation of the auxiliary audio signal. The processing of the generator 105 is therefore subband processing, e.g. a matrix multiplication performed in each subband on the subband samples of the downmix audio signal and the auxiliary audio signal generated by the corresponding subband artificial neural networks.

[0121] The resulting subband / frequency domain representation is then either used directly or converted to a time domain representation using, for example, a suitable synthesis filter bank applied specifically with a separate synthesis filter for each channel.

[0122] Each of the plurality of sub-band neural networks may be referred to as a composite sub-band artificial neural network. The comments given above with respect to the artificial neural network 107 also apply mutatis mutandis to the composite sub-band artificial neural network, and indeed the composite artificial neural network 107 may be considered to be one of the composite sub-band artificial neural networks.

[0123] Thus, in this configuration, each of the synthesis subband artificial neural networks receives subband samples for the subbands of the synthesis subband artificial neural network, and further, all of the synthesis subband artificial neural networks are configured to receive a set of control data values ​​from the source artificial neural network 405.

[0124] Each synthesis subband artificial neural network generates subband samples for a subset of the subbands of the frequency subband representation of the auxiliary audio signal, and typically generates subband samples (only) for the subbands for which it receives input samples for that subband from the filter bank 501.

[0125] In many embodiments, the apparatus includes an artificial neural network for each subband of the frequency subband representation of the auxiliary audio signal generated by the filter bank. Thus, in many embodiments, the output samples for each subband of the filter bank 501 are fed to the input nodes of one synthesis subband artificial neural network, which then generates subband samples of the auxiliary audio signal for that subband. In many embodiments, the subband processing is therefore completely separate for each subband.

[0126] In this example, generation of the auxiliary audio signal is therefore performed on a subband-by-subband basis, with separate and individual artificial neural networks in each subband. The individual artificial neural networks are trained to provide output samples for a subband given the input subband samples for that subband. However, the subband artificial neural networks are further adapted based on control data values ​​received by the receiver 101.

[0127] Such an approach has been found to provide a highly advantageous generation of auxiliary audio signals that enable extremely high-quality reconstruction of multi-channel audio signals. Furthermore, such an approach allows for highly efficient operation with significantly reduced complexity and / or generally significantly reduced computational resource requirements. Subband artificial neural networks tend to be significantly smaller than a single complete artificial neural network required for the generation of the entire signal. Generally, far fewer nodes and, in some cases, fewer layers are required for processing, resulting in a significant reduction in the number of operations and calculations required to perform the artificial neural network function. While more artificial neural networks are required to cover all subbands, smaller artificial neural networks generally significantly reduce the total number of operations required, and therefore the overall computational resource requirements. Furthermore, in many scenarios, this allows for a more efficient training process.

[0128] The subband structure therefore provides a computationally efficient approach for enabling an artificial neural network to be implemented to assist in the decoding of audio data including downmix audio signals and upmix parametric data. The described systems and approaches enable high-quality multi-channel audio signals to be reconstructed, and in general, significantly improved audio quality can be achieved compared to conventional approaches. Furthermore, a computationally efficient decoding process can be achieved. The subband and artificial neural network-based approach is also compatible with other processes that use subband processing.

[0129] In some embodiments, subband processing is more flexible than strict subband-by-subband processing. For example, in some embodiments, each synthesis subband artificial neural network receives not only subband samples from the subband itself, but also, possibly, subband samples for one or more other subbands. For example, a synthesis subband artificial neural network for one subband may, in some embodiments, also receive samples of the downmix audio signal from one or two neighboring / adjacent subbands. As another example, in some embodiments, one or more of the synthesis subband artificial neural networks may also receive input samples from one or more subbands that contain harmonics (or subharmonics) of the frequencies of the subband. For example, a subband around a 500 Hz center frequency may also receive frequencies from a subband around a 1000 Hz center frequency. Such additional subbands, having a specific relationship with the subbands of the synthesis subband artificial neural network, provide additional information that allows for the generation of improved synthesis subband artificial neural networks for some audio signals.

[0130] In some embodiments, all of the composite sub-band artificial neural networks have the same properties and dimensions. In particular, in many embodiments, all of the composite sub-band artificial neural networks have the same number of input and output nodes, and possibly the same internal structure. Such an approach is used, for example, in embodiments in which all of the sub-bands have the same bandwidth.

[0131] In some embodiments, however, the composite sub-band artificial neural network comprises non-equivalent neural networks. In particular, in some embodiments, the number of input nodes for the composite sub-band artificial neural network is different for at least two artificial neural networks. Thus, in some embodiments, the number of input samples included in determining the output samples is different for different sub-band and composite sub-band artificial neural networks.

[0132] In some embodiments, the number of samples / input nodes is greater in some lower frequency subbands than in some higher frequency bands. In fact, the number of samples / input nodes decreases monotonically with increasing frequency. The lower frequency composite subband artificial neural network is therefore larger than the higher frequency composite subband artificial neural network and considers more input samples than the higher frequency composite subband artificial neural network. Such an approach can be combined with subbands having different bandwidths, for example, when the lower frequency subbands have a higher bandwidth than the higher frequency bands.

[0133] Such an approach provides an improved trade-off between achievable audio quality and computational complexity and resource usage in many scenarios, providing for closer adaptation of the system to reflect the general characteristics of audio, thereby enabling more efficient processing.

[0134] In the approach described above, a single source artificial neural network 405 is used to generate control data values ​​for a set of control data values, while multiple synthesis sub-band artificial neural networks process the sub-band downmix audio signal.

[0135] In another embodiment, the generation of the sets of control data values ​​in the source audio device is subband-based. Similar to the synthesis audio device, the source audio device also includes a filter bank 601 applied to the downmix audio signal to provide a set of subband signals, which are then fed to a corresponding subband artificial neural network, each generating a set of control data values, as shown in FIG. 6. Such multiple artificial neural networks are also referred to as source subband artificial neural networks. The sets of control data values ​​generated by the source subband artificial neural network are then fed to the synthesis subband artificial neural network. The subbands generated by such filter bank 601 need not be the same as the subbands generated by filter bank 501 that generates the subband samples for the synthesis subband artificial neural network. For example, each of the subbands for determining the sets of control data values ​​may include a different number of subbands for the synthesis subband artificial neural network, and the set of control data value sets determined by one source subband artificial neural network is fed to the appropriate synthesis subband artificial neural network.

[0136] However, in many embodiments, the subbands used for control data value generation by the source audio device are the same as the subbands used for the receiver's synthesis subband artificial neural network.

[0137] The remarks given above with respect to the source artificial neural network 405 also apply mutatis mutandis to the source sub-band artificial neural network, and indeed the source artificial neural network 405 can be considered to be one of the source sub-band artificial neural networks.

[0138] In this approach, each of the source subband artificial neural networks generates a set of control data values ​​that, in some embodiments, are applied to only a subset of the synthesis subband artificial neural networks. Indeed, in many embodiments, each of the source subband artificial neural networks generates a set of control data values ​​for one of the synthesis subband artificial neural networks. In some embodiments, the apparatus includes the same number of source subband artificial neural networks and synthesis subband artificial neural networks, specifically, one source subband artificial neural network and one synthesis subband artificial neural network for each subband of the filter bank 501. In other embodiments, there are different numbers of source subband artificial neural networks and synthesis subband artificial neural networks. For example, one source subband artificial neural network receives input samples for a group of subbands and generates a set of control data values ​​for that group of subbands. This set of control data values ​​is then applied to the group of synthesis subband artificial neural networks for those subbands.

[0139] Such a subband-based approach for generating a set of control data values ​​provides improved results in many scenarios by enabling a more accurate set of control data values ​​to be generated and used to adapt the synthesis subband artificial neural network. It also reduces complexity and / or resource usage in many scenarios. For example, in many embodiments, fewer values ​​in the set of control data values ​​are achieved, thereby enabling a reduction in the complexity of both the source subband artificial neural network and the synthesis subband artificial neural network. Furthermore, the source subband artificial neural network will generally have fewer inputs than, and be much smaller than, a full-band artificial neural network.

[0140] In many such embodiments, each of the source subband artificial neural networks generates a set of control data values ​​for a given subset of subbands, and generally for a single subband, based on the subband samples of that subset of subbands. However, in addition, one or more of the source subband artificial neural networks further includes subbands that are from one or more other subbands, i.e., the input to the source subband artificial neural network has input nodes that receive subband samples for subbands for which the source subband artificial neural network does not generate any set of control data values.

[0141] As a specific example, in many embodiments, each source subband artificial neural network receives as input not only subsamples from the subband for which it generates a set of control data values, but also subsamples from, for example, adjacent subbands.

[0142] Such an approach often allows for generating an improved set of control data values, which often leads to an improved audio quality. In particular, it has been found that considering surrounding subbands allows the set of control data values ​​to better reflect the temporal resolution of the downmix audio signal / multi-channel audio signal. It has been found that such an approach in particular allows for a better representation of the temporal kurtosis.

[0143] In some embodiments, one or more of the source subband artificial neural networks, or indeed the source artificial neural network 405 if only one such artificial neural network is included, further includes as input subband samples from outside the time interval in which the corresponding synthesis subband artificial neural network generates samples of the auxiliary audio signal.

[0144] In particular, the audio device processing operates on a frame-by-frame basis, where a time interval / frame of the received downmix audio signal is processed to generate output samples for the multi-channel audio signal for that time interval / frame. Thus, for each frame, the filter bank 501 generates subband samples which are fed to a source subband artificial neural network which generates a set of control data values ​​for the time interval, and to a source subband artificial neural network which generates subband samples of a subband representation of the multi-channel audio signal based on the subbands and the set of control data values.

[0145] Thus, in particular, each source subband artificial neural network operates in a block manner, each operation in which a set of output samples is generated from a set of input samples corresponding to a time interval of the downmix audio signal / multi-channel audio signal for which an output sample of the multi-channel audio signal is generated.

[0146] In some embodiments, one or more of the source subband artificial neural networks, in addition to the appropriate subband samples generated for the current time interval, also receive subband samples for another time interval, such as typically from one or more adjacent time intervals. For example, in some embodiments, one or more of the source subband artificial neural networks also includes subband samples for the previous and next time intervals.

[0147] In many embodiments, such an approach provides for an improved set of control data values ​​to be generated, which leads to improved audio quality.

[0148] In some embodiments, at least some of the sets of control data values ​​are common for multiple subbands of the frequency subband representation of the downmix audio signal.

[0149] In the above-described approaches, all of the composite sub-band artificial neural networks are provided with the same set of control data values. In some embodiments, only some of the composite sub-band artificial neural networks are provided with the same set of control data values. In particular, in some embodiments, at least one control data value of the set of control data values ​​is processed by at least two composite sub-band artificial neural networks.

[0150] By using the same set of control data values, many embodiments achieve improved efficiency and performance, often reducing complexity and resource usage in generating the set of control data values. Furthermore, in many scenarios, it provides improved performance, as all available information provided by the control data values ​​is considered by each synthesis sub-band artificial neural network, thus achieving improved adaptation of the synthesis sub-band artificial neural network.

[0151] However, in some embodiments, different synthesis sub-band artificial neural networks are provided with different sets of control data values, and in particular, in some embodiments, one or more control data values ​​are processed by one of the synthesis sub-band artificial neural networks but not by another of the synthesis sub-band artificial neural networks.

[0152] For example, the source artificial neural network 405 generates a set of control data values, and different subsets of these are provided to different synthesis sub-band artificial neural networks. In other embodiments, several control data sets are also generated, including several control data values ​​that are provided manually or generated, for example, by analysis of the downmix audio signal. For example, harmonics or peaks are detected in the downmix audio signal. Such data are, for example, applied to only some of the synthesis sub-band artificial neural networks. For example, detected peaks or harmonics are presented only to the synthesis sub-band artificial neural networks of the sub-bands in which they were detected.

[0153] In many embodiments, different composite sub-band artificial neural networks are thus provided with different sets of control data values, which often involves some control data values ​​being the same and some control data values ​​being different for different sets of composite sub-band artificial neural networks.

[0154] The above description has focused on the application of subband processing to a downmix audio signal and to a source subband artificial neural network 107 that generates a set of control data values ​​based on the downmix audio signal. However, it will be understood that the described techniques and principles may be equally applied to multi-channel audio signals, and in fact, with minor modifications, references to a downmix audio signal may be considered to be equivalent to references to a multi-channel audio signal. Thus, in particular, subband processing is applied to a multi-channel audio signal by applying a filter bank to the channels of the multi-channel audio signal.

[0155] Artificial neural networks are adapted to specific purposes through a training process used to adapt / tune / change the weights and other parameters (e.g., biases) of the artificial neural network. It will be appreciated that many different training processes and algorithms for training artificial neural networks are known. Typically, training is based on a large training set, in which a large number of examples of input data are presented to the network. Furthermore, the output of the artificial neural network is typically compared (directly or indirectly) to expected or ideal results. A cost function is generated to reflect the desired outcome of the training process. In a common scenario known as supervised learning, the cost function often represents the distance between the prediction for specific input data and the ground truth. Based on the cost function, the weights are modified, and by repeating the process with the modified weights, the artificial neural network is adapted toward a state where the cost function is minimized.

[0156] More specifically, during the training step, a neural network has two distinct flows of information: from input to output (forward pass) and from output to input (backward pass). In the forward pass, data is processed by the neural network as described above, while in the backward pass, weights are updated to minimize a cost function. Generally, such backward propagation follows the gradient direction of the cost function landscape. In other words, by comparing the predicted output with the ground truth for a batch of data input, the direction in which the cost function is minimized and propagated backward can be estimated by updating the weights accordingly. Other known techniques for training artificial neural networks include, for example, the Levenberg-Marquardt algorithm, the conjugate gradient method, and Newton's method.

[0157] In this case, training specifically involves a training set containing a potentially large number of multi-channel audio signals. In some embodiments, the training data are multi-channel audio signals in time segments corresponding to the processing time interval of the artificial neural network being trained, e.g., the number of samples in the training multi-channel audio signals corresponds to the number of samples corresponding to the input nodes of the artificial neural network being trained. Each training example therefore corresponds to one operation of the artificial neural network being trained. Typically, however, batches of training samples are considered for each step to accelerate the training process. Furthermore, many improvements to gradient descent also make it possible to accelerate convergence or avoid local minima in the cost function landscape.

[0158] For each training multi-channel audio signal, the training processor performs a downmix operation to generate a downmix audio signal and corresponding upmix parametric data. Thus, the encoding process applied to the multi-channel audio signals during normal operation is also applied to the training multi-channel audio signals, thereby generating the downmix and upmix parametric data.

[0159] Furthermore, the training processor, in some embodiments, generates a residual signal that reflects differences between the downmix audio signal and the multi-channel audio signal, or more generally, that represents parts of the multi-channel audio signal that are not adequately represented by the downmix audio signal. For example, in many embodiments, the training processor generates a downmix signal and also generates a residual signal that, when used in upmixing based on the upmix parametric data, results in a (more) accurate reconstructed multi-channel audio signal.

[0160] In particular, for stereo multi-channel audio signals, the training processor uses a parametric stereo scheme (e.g., according to a suitable standardized method). Such encoding applies frequency- and time-dependent matrix operations, such as rotation operations, to the input stereo signal to generate a downmix signal and a residual signal. For example, a 2×2 matrix / complex-valued multiplication is typically applied to the input stereo signal to substantially align one of the rotated channel signals to have the maximum signal value. This channel is used as a mono signal, and the rotation is typically performed frame-by-frame. The rotation value is stored as part of the upmix parametric data (or a parameter that allows this determination is included in the upmix parametric data). Thus, in the synthesizer, an inverse rotation is performed to reconstruct the stereo signal. Rotating the stereo signal results in another stereo signal, one of whose channels is accordingly aligned with maximum intensity. The other channels are typically discarded in the parametric stereo encoder to reduce the data rate. In conventional PS decoding, a decorrelated signal is typically generated in the decoder and used for the upmixing process. In current training techniques, this second signal is used as the residual signal for downmixing because it represents information discarded in the encoder, and therefore it represents the ideal signal to be reconstructed in the decoder as part of the upmixing process.

[0161] Thus, in some embodiments, the training processor generates a training downmix signal and / or a training residual signal from the training multi-channel audio signal. The training downmix signal is fed to an arrangement of a source artificial neural network 405 and a synthesis artificial neural network 107, or equivalently to a source sub-band artificial neural network and a synthesis sub-band artificial neural network arrangement, i.e. samples of the training downmix audio signal are fed to a neural network that uses the same processing as that applied to the downmix audio signal by the audio device during normal operation (including, for example, sub-band filtering, etc.).

[0162] Thus, in this approach, joint training of the source (sub-band) and synthesis (sub-band) artificial neural networks is applied, and the input samples supplied to the artificial neural networks in the source and synthesis audio devices are generated for the training multi-channel audio signal in the same way as for the real signals when the devices are in normal operation. The artificial neural networks of the source and synthesis audio devices are thus trained together, and the weights etc. of both artificial neural networks are adapted together. If only one of the devices is available, a training configuration is used, in which the missing artificial neural network is emulated / copied by a dedicated training artificial neural network.

[0163] For example, if the source audio device is trained independently from the synthesis audio device, the training set includes standard or generic synthesis audio device features (e.g., it includes an artificial neural network corresponding to the synthesis artificial neural network). In such cases, references below to synthesis audio device features will be considered to refer to the corresponding features in the training set.

[0164] Similarly, if the synthesized audio device is trained independently of any source audio device, the training set includes standard or generic source audio device features (e.g., it includes an artificial neural network that corresponds to the source artificial neural network). In such cases, references below to source audio device features will be considered to refer to the corresponding features in the training set.

[0165] An output from the neural network operation is then determined, and a cost function is applied to determine a cost value for each training downmix audio signal and / or for the combined set of training downmix audio signals (e.g., an average cost value for the training set is determined). The cost function includes various components.

[0166] Generally, the cost function includes at least one component that reflects how close the generated signal is to the reference signal, i.e., the so-called reconstruction error. In some embodiments, the cost function includes at least one component that reflects how close the generated signal is to the reference signal from a perceptual point of view.

[0167] For example, in some embodiments, the auxiliary audio signal generated by the synthesis artificial neural network 107 for a given training downmix audio signal / multi-channel audio signal is compared with the residual signal for that training downmix audio signal / multi-channel audio signal. A cost function contribution / combination is generated that reflects the difference between the generated auxiliary audio signal and the reference residual signal. This process is performed for all training sets to generate the entire cost function.

[0168] Such an example is shown in Figure 7. In this example, a downmixer 701 receives a training multi-channel audio signal, which in a particular example is a stereo signal. For a given training signal, the downmixer 701 performs a downmixing operation to generate a training downmix signal, which in a particular example is a training mono audio signal. The downmix audio signal is fed to a preprocessor 703, which generates, in particular, downmix audio signal samples to be input to the artificial neural network. The preprocessor 703 performs the same operations as those performed in the source audio device and / or synthesis device, in particular to generate samples for the artificial neural network; in fact, the same functions are generally used, i.e., the functions for generating the artificial neural network input samples are also used during the training process.

[0169] The output of the pre-processor 703 is supplied to the source artificial neural network 405 and the synthesis artificial neural network 107, as in the audio device of Figure 1. The source artificial neural network 405 and the synthesis artificial neural network 107 are coupled to each other such that an output set of control data values ​​generated by the source artificial neural network 405 is supplied to the synthesis artificial neural network 107. The output of the synthesis artificial neural network 107 therefore corresponds to samples of the auxiliary audio signal generated for the training downmix audio signal.

[0170] Thus, for a given training multi-channel audio signal, the training system of FIG. 7 performs the same operations as the audio devices of FIGS. 1 and 4, thereby generating the auxiliary audio signal that would be generated by the audio device if the artificial neural network had the same data / configuration (coefficients, biases, etc.).

[0171] The output of the synthetic artificial neural network 107 is fed to a comparator 705, which proceeds to compare the generated auxiliary audio signal with the residual signal generated by the downmixer 701. As mentioned above, the residual signal allows, in principle, a substantially perfect reconstruction of a multi-channel audio signal and can therefore be considered a close approximation of the ideal auxiliary audio signal. The comparison between the generated auxiliary audio signal and the residual signal therefore gives an indication of how advantageous the generated auxiliary audio signal is. A cost value is therefore determined based on the comparison; in particular, the greater the difference, the greater the cost value.

[0172] It will be appreciated that many different techniques can be used to determine a cost value that reflects the difference between the signals. For example, correlation is performed using a cost value that monotonically decreases as the correlation value increases. As another example, two signals are subtracted from each other and a power measure for the difference signal is used as the cost value. It will be appreciated that many other techniques are available and can be used.

[0173] Thus, in this example, the cost function generates a cost value that reflects how closely the generated auxiliary audio signal matches the corresponding residual signal for the training multi-channel audio signal.

[0174] Based on the cost values, the training processor 707 adapts the weights of the artificial neural networks. For example, a backpropagation technique is used. In particular, the training processor 707 adjusts the weights of both the source artificial neural network 405 and the composite artificial neural network 107 based on the cost values. For example, given the derivative (representing the slope) of the weights with respect to the cost function, the weight values ​​are changed to proceed in the direction of the gradient. For simple / minimal computation, the case of a single data input backward pass can be referred to as training a perceptron (single neuron).

[0175] This process is repeated until the artificial neural network is considered trained. For example, training is performed for a predetermined number of iterations. As another example, training continues until the weights change less than a predetermined amount. Also very commonly, a validation step is performed in which the network is tested against a validation metric and stopped when it reaches an expected result.

[0176] As a concrete example, a stereo signal is fed to a traditional PS downmix module, which generates both a downmix and an ideal residual signal, i.e., a residual signal that allows for (near) perfect reconstruction of the waveform at the decoder side. Using the mono signal and residual signal pair, the source artificial neural network 405 generates a set of control data values ​​that give a latent representation of the audio signal (frame). This set of control data values, together with the mono audio signal, is fed to the synthesis artificial neural network 107, which generates the auxiliary audio signal. When the artificial neural network is deep enough, this structure is trained using a cost function, for example, RMSE (root mean square error for the auxiliary audio signal relative to the residual signal), so that the artificial neural network structure learns what the auxiliary audio signal should be like for a given mono audio signal.

[0177] In this configuration, the source artificial neural network 405 and the synthesis artificial neural network 107 are therefore jointly trained using the same downmix audio signal and cost function, and the weights of both the source artificial neural network 405 and the synthesis artificial neural network 107 are updated based on the same training data, the same cost function, and the downmix audio signal.

[0178] In many embodiments, the residual signal is used directly when comparing with the generated auxiliary audio signal. However, more commonly, a target audio signal is generated from the residual audio signal and used in comparison with the generated auxiliary audio signal. The target audio signal is generated by applying a function or signal processing application to the residual signal. For example, scaling / level setting is applied to the residual audio signal to generate the target audio signal. As another example, a filter operation is applied to the residual audio signal to generate the target audio signal. As another example, scaling is applied to the residual signal to minimize differences and / or maximize correlation with the generated auxiliary audio signal.

[0179] In some embodiments, the cost function is alternatively or additionally configured to reflect the difference between the training multi-channel audio signal and the multi-channel audio signal generated by upmixing the downmix audio signal and the generated auxiliary audio signal. In such examples, the downmixer 701 also generates upmix parametric data used in the upmixing. Thus, in some embodiments, rather than simply training the artificial neural network to generate an auxiliary audio signal that matches the residual signal, the training includes a reconstruction of the multi-channel audio signal based on the generated auxiliary audio signal. The generation of the output multi-channel audio signal is performed, in particular, using the same operations as those performed in an encoder. This output multi-channel audio signal is then compared with the input multi-channel audio signal. Thus, in some embodiments, the training data does not essentially include the training downmix audio signal, but instead, the training data alternatively or additionally includes the multi-channel audio signal. The cost function may also be based on a comparison of the generated auxiliary audio signal with, for example, the residual signal, or, for example, between the original training multi-channel audio signal and the reconstructed multi-channel audio signal.

[0180] In some embodiments, the cost function further considers other parameters. For example, in some embodiments, the cost function further includes consideration of the degree of correlation between the generated auxiliary audio signal and the downmix signal. In particular, the cost function indicates that the lower the cost, the more the downmix audio signal and the auxiliary audio signal are decorrelated. Increased decorrelation indicates a decrease in the common information present in both of the two signals.

[0181] In embodiments where subband operations are performed, the described techniques are performed for each subband. In particular, a residual audio signal is generated for each subband and compared with the generated auxiliary audio signal subband samples for that subband. Accordingly, a cost function is evaluated for each subband to determine a cost value, and the coefficients of the source subband artificial neural network and the synthesis subband artificial neural network for that subband are trained.

[0182] In many embodiments, the source subband artificial neural networks, and / or indeed the source artificial neural network 405, generate sets of control data values ​​that match the time and frequency resolution of the composite subband artificial neural network, including, in particular, the composite artificial neural network 107 when only one artificial neural network is used. However, in other embodiments, the time and / or frequency resolution of the generated sets of control data values ​​is different, typically lower for those sets of control data values. In such situations, functionality for changing the resolution is included. In particular, an interpolator is incorporated to generate an interpolated set of parameter values ​​from the generated parameter values.

[0183] The audio device is in particular implemented in one or more suitably programmed processors. In particular, the artificial neural network is implemented in one or more such suitably programmed processors. The different functional blocks, in particular the artificial neural network, may be implemented in separate processors and / or may, for example, be implemented in the same processor. An example of a suitable processor is provided below.

[0184] 8 is a block diagram illustrating an exemplary processor 800, according to an embodiment of the present disclosure. Processor 800 may be used to implement one or more processors that implement the previously described devices or elements thereof (including, inter alia, one or more artificial neural networks). Processor 800 may be any suitable processor type, including, but not limited to, a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable array (FPGA) programmed to form a processor, an FPGA, a graphic processing unit (GPU), an application specific integrated circuit (ASIC) designed to form a processor, or a combination thereof.

[0185] Processor 800 includes one or more cores 802. Core 802 includes one or more arithmetic logic units (ALUs) 804. In some embodiments, core 802 includes a floating point logic unit (FPLU) 806 and / or a digital signal processing unit (DSPU) 808 in addition to or instead of the ALUs 804.

[0186] The processor 800 includes one or more registers 812 communicatively coupled to the core 802. The registers 812 are implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 812 are implemented using static memory. The registers provide data, instructions, and addresses to the core 802.

[0187] In some embodiments, processor 800 includes one or more levels of cache memory 810 communicatively coupled to cores 802. Cache memory 810 provides computer-readable instructions to cores 802 for execution. Cache memory 810 provides data for processing by cores 802. In some embodiments, the computer-readable instructions are provided to cache memory 810 by local memory, for example, local memory attached to external bus 816. Cache memory 810 may be implemented using any suitable cache memory type, for example, metal-oxide-semiconductor (MOS) memory such as static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology.

[0188] Processor 800 includes a controller 814, which controls input to processor 800 from other processors and / or components included in the system and / or output from processor 800 to other processors and / or components included in the system. Controller 814 controls data paths in ALU 804, FPLU 806, and / or DSPU 808. Controller 814 is implemented as one or more state machines, data paths, and / or dedicated control logic. Gates in controller 814 are implemented as standalone gates, FPGAs, ASICs, or any other suitable technology.

[0189] Registers 812 and cache 810 communicate with controller 814 and core 802 via internal connections 820A, 820B, 820C, and 820D. The internal connections may be implemented as buses, multiplexers, crossbar switches, and / or any other suitable connection technology.

[0190] Input and output for processor 800 are provided via bus 816, which may include one or more conductive lines. Bus 816 is communicatively coupled to one or more components of processor 800, such as controller 814, cache 810, and / or registers 812. Bus 816 is coupled to one or more components of the system.

[0191] Bus 816 is coupled to one or more external memories. The external memory includes read-only memory (ROM) 832. ROM 832 may be masked ROM, electronically erasable programmable read-only memory (EPROM), or any other suitable technology. The external memory includes random access memory (RAM) 833. RAM 833 may be static RAM, battery-backed static RAM, dynamic RAM (DRAM), or any other suitable technology. The external memory includes electrically erasable programmable read-only memory (EEPROM®) 835. The external memory includes flash memory 834. The external memory includes a magnetic storage device, such as a disk 836. In some embodiments, the external memory is included within the system.

[0192] The invention may be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention is optionally implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of an embodiment of the invention may be physically, functionally, and logically implemented in any suitable way. Indeed, functionality may be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the invention may be implemented in a single unit, or may be physically and functionally distributed between different units, circuits, and processors.

[0193] While the present invention has been described with reference to several embodiments, it is not intended that the present invention be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Furthermore, while features may appear to be described with reference to particular embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the terms "comprise" or "include" do not exclude the presence of other elements or steps.

[0194] Furthermore, although individually listed, a plurality of means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Moreover, although individual features may be included in different claims, these may, in some cases, be advantageously combined, and their inclusion in different claims does not imply that the combination of features is infeasible and / or unadvantageous. Furthermore, the inclusion of a feature in one category of claims does not imply limitation to this category, but rather indicates that the feature is equally applicable to other claim categories, where appropriate. Furthermore, the order of features in the claims does not imply that those features must be performed in any particular order, and in particular the order of individual steps in method claims does not imply that those steps must be performed in that order. Rather, steps may be performed in any suitable order. Furthermore, singular references do not exclude pluralities; thus, references to "first," "second," etc. do not exclude pluralities. Reference signs in the claims are provided merely as an illustrative example and shall not be construed in any way as limiting the scope of the claims.

Claims

1. A device for generating multi-channel audio signals, A downmixed audio signal for the aforementioned multi-channel audio signal, Upmix parametric data for upmixing the downmix audio signal, A set of control data values ​​that adapts a process for generating an auxiliary signal for upmixing from the downmix audio signal. A receiver for receiving data signals, An artificial neural network having an input node for receiving samples of the downmix audio signal and an output node for supplying samples of an auxiliary audio signal, further comprising a node provided for including contributions from the set of control data values, A generator for generating the multi-channel audio signal from the downmix audio signal and the auxiliary signal, depending on the upmix parametric data. A device equipped with the following features.

2. The apparatus according to claim 1, further comprising a filter bank for generating frequency subband representations of the downmixed audio signal, wherein the sample of the downmixed audio signal is a subband sample of the frequency subband representation.

3. The apparatus according to claim 2, wherein the artificial neural network is an artificial neural network among a first plurality of subband artificial neural networks, and each subband artificial neural network of the first plurality of subband artificial neural networks generates subband samples for a subset of subbands of the frequency subband representation of the auxiliary audio signal.

4. The apparatus according to claim 3, wherein at least one control data value from the set of control data values ​​is processed by at least two artificial neural networks from the first plurality of subband artificial neural networks.

5. The apparatus according to claim 3, wherein at least one of the set of control data values ​​is not processed by at least one of the first plurality of subband artificial neural networks.

6. The apparatus according to claim 1, wherein the artificial neural network is trained with training data comprising at least one of a training multichannel audio signal, a training downmix audio signal, and training upmix parametric data generated from the training multichannel audio signal, and the training employs a cost function that compares the training multichannel audio signal with an upmixed multichannel signal generated from the training downmix audio signal and the generated auxiliary audio signal using the training upmix parametric data.

7. The apparatus according to any one of claims 1 to 6, wherein the artificial neural network is trained by training data having training input data including a training downmix audio signal of a training multichannel audio signal and a training set of control data values ​​for the training downmix audio signal, and using a cost function that includes a contribution indicating the difference between a training auxiliary audio signal generated by the artificial neural network in response to the training data and a training residual signal for the training downmix audio signal.

8. The apparatus according to claim 7, wherein the cost function further includes a contribution indicating the degree of correlation between the auxiliary audio signal and the downmixed audio signal of the multichannel audio signal.

9. The apparatus according to claim 1, further comprising an upsampler for temporally upsampling the set of control data values, wherein the artificial neural network includes contributions from the upsampled control data values.

10. A device for generating data signals that represent multi-channel audio signals, wherein the device is A downmixer for generating a downmix audio signal by downmixing the aforementioned multichannel audio signal, further comprising a downmixer for generating upmix parametric data for upmixing the downmix audio signal, An artificial neural network for generating a set of control data values ​​that adapts processing for generating an auxiliary audio signal for upmixing from the downmix audio signal, the artificial neural network having an output node for supplying the set of control data values ​​and an input node for receiving a sample of at least one of the downmix audio signal and the multichannel audio signal, A generator for generating the data signal, which includes the downmix audio signal, the upmix parametric data, and the set of control data values. A device equipped with the following features.

11. The apparatus according to claim 10, wherein the set of control data values ​​provides a latent representation of at least one of the downmix audio signal and the multichannel audio signal.

12. The apparatus according to claim 10 or 11, wherein the artificial neural network is trained by training data having training input data which is a sample of a training audio signal, which is at least one of a multichannel audio signal and a downmix audio signal of the multichannel audio signal, and by using a cost function which includes a component indicating the difference between the training audio signal and an output audio signal generated from a set of control data values ​​generated by the artificial neural network for the training audio signal.

13. The apparatus according to claim 12, wherein the output audio signal is the output of a synthetic artificial neural network having an input node that receives the set of control data values ​​generated by the artificial neural network for the training audio signal, and the artificial neural network and the synthetic artificial neural network are trained together.

14. The apparatus according to claim 12, wherein the training audio signal and the output audio signal are downmix audio channels.

15. A method for generating a multi-channel audio signal, A downmixed audio signal for the aforementioned multi-channel audio signal, Upmix parametric data for upmixing the downmix audio signal, A set of control data values ​​that adapts a process for generating an auxiliary signal for upmixing from the downmix audio signal. A step of receiving a data signal, including, The artificial neural network has an input node that receives samples of the downmix audio signal and an output node that supplies samples of an auxiliary audio signal, and further comprises a node provided for including contributions from the set of control data values, The steps include generating the multi-channel audio signal from the downmix audio signal and the auxiliary signal, depending on the upmix parametric data, and A method having.

16. A method for generating a data signal that represents a multi-channel audio signal, A step of generating a downmix audio signal by downmixing the multichannel audio signal, wherein the downmixer further generates upmix parametric data for upmixing the downmix audio signal, Steps include: an artificial neural network generating a set of control data values ​​that adapt processing for generating an auxiliary audio signal for upmixing from the downmix audio signal; the artificial neural network having an output node that provides the set of control data values ​​and an input node that receives a sample of at least one of the downmix audio signal and the multichannel audio signal; A step of generating the data signal to include the downmix audio signal, the upmix parametric data, and the set of control data values. A method having.

17. A computer program comprising computer program code means for performing all steps of the method according to claim 15 or 16 when run on a computer.