Multi-channel audio signal generation

JP2025529995A5Pending Publication Date: 2026-08-26KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025514545
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-13
Filing Date
2023-08-29
Publication Date
2026-08-26

AI Technical Summary

Technical Problem

Existing parametric multi-channel audio coding techniques introduce distortions, shifts, and artifacts, leading to reduced audio quality and higher data rates, while also increasing processing complexity and resource usage.

Method used

Employing trained artificial neural networks to generate multi-channel audio signals from a downmix signal and upmix parametric data, utilizing a first neural network to extract features and a second neural network to generate an auxiliary audio signal, which is combined with the downmix signal to produce a high-quality multi-channel output through matrix multiplication.

Benefits of technology

This approach enhances audio quality, reduces complexity and resource usage, and improves the data rate tradeoff by providing efficient and flexible multi-channel audio signal generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The audio device includes a receiver 101 that receives a downmix audio signal of a multi-channel audio signal and upmix parametric data for upmixing the downmix audio signal. A first artificial neural network 107 generates a set of features of the downmix audio signal from samples of the downmix audio signal. A second artificial neural network 109 has an input node that receives a second sample of the downmix audio signal and a node that receives a feature from the set of features. Based on these inputs, the second artificial neural network 109 generates samples of an auxiliary audio signal of the downmix audio signal. A generator 105 generates a multi-channel audio signal from the downmix signal and the auxiliary audio signal depending on the upmix parametric data. In many embodiments, the calculations are subband-based, and separate artificial neural networks are used for different subbands.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the generation of multi-channel audio signals, and in particular, but not exclusively, to the generation of a stereo signal from upmixing of a mono downmix signal using upmix parametric data. [Background technology]

[0002] Spatial audio applications are becoming numerous and widespread, increasingly forming at least a part of many audiovisual experiences. Indeed, new and improved spatial experiences and applications are continually being developed, resulting in increased demands on audio processing and rendering.

[0003] For example, interest in virtual reality (VR) and augmented reality (AR) has grown in recent years, with many implementations and applications reaching the consumer market. Indeed, equipment is being developed to both render the experience and capture or record data suitable for such applications. For example, relatively low-cost equipment is being developed to enable game consoles to provide full VR experiences. With the VR and AR markets expected to reach significant size in the short term, this trend is expected to continue and certainly gain pace. In the audio field, there is a prominent field of research devoted to the reproduction and synthesis of realistic and natural spatial audio. The ideal goal is to create audio sources so natural that users cannot tell the difference between the synthesized and the original.

[0004] Many research and development efforts have focused on providing efficient, high-quality audio coding and decoding for spatial audio. Multi-channel audio representations, including stereo representations, are often used as spatial audio representations, and efficient coding of multi-channel audio based on downmixing a multi-channel audio signal to a downmix channel with fewer channels has been developed. One of the major advances in low-bitrate audio coding is the use of parametric multi-channel coding, in which a downmix signal is generated together with parametric data, which can be used to upmix the downmix signal to recreate the multi-channel audio signal.

[0005] Specifically, parametric multi-channel audio coding, instead of traditional mid-side coding or intensity coding, downmixes a multi-channel input signal to a smaller number of channels (e.g., 2:1) and extracts multi-channel image (stereo) parameters. The downmix signal is then encoded using a more traditional audio coder (e.g., a mono audio encoder). The downmix bitstream is multiplexed with the coded multi-channel image parameter bitstream. This bitstream is then sent to a decoder, where the process is reversed. First, the downmix audio signal is decoded, and then the multi-channel audio signal is reconstructed guided by the coded multi-channel image upmix parameters.

[0006] An example of stereo coding is described in E. Schuijers, W. Oomen, B. den Brinker, and J. Breebaart, "Advances in Parametric Coding for High-Quality Audio," 114th AES Convention, Amsterdam, the Netherlands, 2003, Preprint 5852. In the described technique, the downmixed mono signal is parameterized by exploiting the natural separation of a signal into three components (objects): transients, sinusoids, and noise. E. Schuijers, J. Breebaart, H. Pumhagen, and J. Engdegard, "Low Complexity Parametric Stereo Coding," 116th AES Convention, Berlin, Germany, 2004, Preprint 6073, provides details on how to achieve low (decoder) complexity when combining parametric stereo with spectral band replication (SBR).

[0007] In the described approach, decoding is based on the use of a so-called decorrelation process, in which a decorrelated helper signal is generated from the mono signal. In the stereo reconstruction process, both the mono signal and the decorrelated helper signal are used to generate an upmixed stereo signal based on the upmix parameters. Specifically, the two signals are multiplied by a time- and frequency-dependent 2x2 matrix with coefficients determined from the upmix parameters to provide the output stereo signal.

[0008] However, while parametric stereo (PS) and similar downmix encoding / decoding techniques represent a departure from traditional stereo and multi-channel coding, they are not optimal in all scenarios. In particular, known encoding and decoding techniques tend to introduce distortions, shifts, and artifacts that can result in differences between the (original) multi-channel audio signal input to the encoder and the reproduced multi-channel audio signal at the decoder. This typically results in reduced audio quality and imperfect multi-channel reproduction. Furthermore, data rates remain higher than desired, and processing complexity / resource usage is higher than recommended. Summary of the Invention [Problem to be solved by the invention]

[0009] As such, improved techniques would be advantageous, particularly techniques that allow for increased flexibility, improved adaptability, improved performance, improved audio quality, improved audio quality to data rate tradeoff, reduced complexity and / or resource usage, reduced computational burden, easier implementation, and / or an improved spatial audio experience.

[0010] SUMMARY OF THE INVENTION Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination. [Means for solving the problem]

[0011] According to one aspect of the present invention, there is provided an apparatus for generating a multi-channel audio signal, the apparatus comprising: a receiver for receiving a downmix audio signal of the multi-channel audio signal and upmix parametric data for upmixing the downmix audio signal, a first artificial neural network for generating a set of features of the downmix audio signal, the first artificial neural network having an input node for receiving a first sample of the downmix audio signal and an output node for providing the set of features, a second artificial neural network having an input node for receiving a second sample of the downmix audio signal and an output node for providing a sample of an auxiliary audio signal for the downmix audio signal, the second artificial neural network further comprising a node for receiving a feature from the set of features, and a generator for generating the multi-channel audio signal from the downmix signal and the auxiliary audio signal depending on the upmix parametric data.

[0012] This approach, in many embodiments, improves the audio experience. For many signals and scenarios, this approach improves the generation / reconstruction of multi-channel audio signals with improved perceived audio quality. In many embodiments and scenarios, this approach provides a particularly advantageous arrangement that allows for facilitating and / or facilitating the use of artificial neural networks in audio processing, typically including audio encoding and / or decoding. This approach allows for the advantageous use of artificial neural networks in generating multi-channel audio signals from downmix audio signals.

[0013] This approach provides an efficient implementation and, in many embodiments, can reduce complexity and / or resource usage, and in many scenarios allows the downmix signal to be used to reduce the data rate of data representing a multi-channel audio signal.

[0014] The first and second samples may be the same or different (or partially the same) samples. The first and second samples may be time-domain samples, frequency-domain samples, or may span a particular time and frequency range (particularly sub-band domain samples). The samples of the auxiliary audio signal may be time-domain samples, frequency-domain samples, or may span a particular time and frequency range (particularly sub-band domain samples).

[0015] The upmix parametric data may include parameters (values) relating properties of the downmix signal to properties of the multi-channel audio signal. The upmix parametric data may include data indicating relative properties between channels of the multi-channel audio signal. The upmix parametric data may include data indicating differences in properties between channels of the multi-channel audio signal. The upmix parametric data may include data perceptually related to the synthesis of the multi-channel audio signal. Properties are, for example, differences in phase, magnitude, timing, and / or correlation. In some embodiments and scenarios, the upmix parametric data may represent abstract properties that are not immediately understandable to humans / experts (but typically facilitate better reconstruction / data rate reduction, etc.). The upmix parametric data may include data including at least one of inter-channel magnitude difference, inter-channel timing difference, inter-channel correlation, and / or inter-channel phase difference for the channels of the multi-channel audio signal.

[0016] The first and second artificial neural networks are trained artificial neural networks.

[0017] The first and / or second artificial neural networks are artificial neural networks trained with training data including a training downmix audio signal and training upmix parametric data generated from the training multi-channel audio signal. The training uses a cost function that compares the training multi-channel audio signal with an upmixed multi-channel signal generated from the training downmix signal and the generated auxiliary audio signal using the training upmix parametric data. The first and / or second artificial neural networks may be trained artificial neural networks trained with training data including training data representing various relevant audio sources, including video, film, telecommunication, etc. recordings.

[0018] The first and / or second artificial neural network may be a trained artificial neural network that is trained with training data having training input data including a training downmix audio signal of the training multi-channel audio signal and using a cost function that includes a contribution indicative of a difference between a training auxiliary audio signal generated by the second artificial neural network in response to the training data and a training residual signal of the training downmix audio signal.

[0019] The generator is able to generate a multi-channel audio signal by applying a matrix multiplication to the downmix signal and the auxiliary audio signal, where the coefficients of this matrix are determined as a function of the parameters of the upmix parametric data, and the matrix is ​​time and frequency dependent.

[0020] The audio device is specifically an audio decoder device.

[0021] According to an optional feature of the invention, the apparatus includes a first filter bank for generating a frequency subband representation of the downmix audio signal, and at least some of the second samples of the downmix audio signal are subband samples of the frequency subband representation.

[0022] Sub-band processing provides particularly advantageous computations in many embodiments, and the present arrangement is particularly suited to sub-band processing, resulting in reduced complexity and / or improved multi-channel audio signals.

[0023] According to an optional feature of the invention, the second artificial neural network is an artificial neural network consisting of a first plurality of subband artificial neural networks, each subband artificial neural network of the first plurality of subband artificial neural networks generating subband samples for a subset of subbands of the frequency subband representation of the auxiliary audio signal.

[0024] A particular advantage of this approach is that it allows for highly efficient sub-band processing, thereby splitting the required processing into multiple smaller artificial neural networks, which typically results in reduced complexity and / or improved multi-channel audio signals.

[0025] In many embodiments, each (or at least some) subband neural network generates subband samples for one subband of the frequency subband representation of the auxiliary audio signal.

[0026] According to an optional feature of the invention, the plurality of subband artificial neural networks includes an artificial neural network for each subband of the frequency subband representation of the auxiliary audio signal.

[0027] This provides very advantageous and efficient implementation, computation, and / or performance in many embodiments and scenarios.

[0028] According to an optional feature of the invention, the generator generates the frequency subband representations of the multi-channel audio signal by applying a subband matrix operation to the frequency subband representations of the auxiliary audio signal and the frequency subband representations of the downmix audio signal, and converts the frequency subband representations of the multi-channel audio signal into a time domain representation of the multi-channel audio signal.

[0029] This provides very advantageous and efficient implementation, computation, and / or performance in many embodiments and scenarios.

[0030] According to an optional feature of the invention, a set of features generated by a subband artificial neural network of the first plurality of subband artificial neural networks is common to a plurality of subbands of the frequency subband representation of the downmix audio signal.

[0031] This results in a particularly efficient implementation and / or improved performance.

[0032] The set of features generated by the subband neural networks and common to the plurality of subbands is input to a plurality of subband artificial neural networks (in particular, the second artificial neural network is one of the plurality of artificial neural networks having nodes that receive the common set of features) that generate the multi-channel audio signal.

[0033] According to an optional feature of the invention, the number of input nodes of the artificial neural networks of the first plurality of subband artificial neural networks decreases monotonically with increasing frequency.

[0034] This results in a particularly efficient implementation and / or improved performance.

[0035] According to an optional feature of the invention, the apparatus includes a second filter bank for generating a frequency subband representation of the downmix audio signal, at least some of the first samples of the downmix audio signal being subband samples of the frequency subband representation.

[0036] This results in a particularly efficient implementation and / or improved performance.

[0037] Sub-band processing provides particularly advantageous computations in many embodiments, and the present arrangement is particularly suited to sub-band processing, resulting in reduced complexity and / or improved multi-channel audio signals.

[0038] The first filter bank and the second filter bank may be the same or different, and the subband representation of the downmix audio signal generated by the first filter bank and fed to the first plurality of subbands may use the same subbands as the subband representation of the downmix audio signal generated by the second filter bank and fed to the second plurality of subbands, or in some embodiments may be different.

[0039] According to an optional feature of the invention, the first artificial neural network is an artificial neural network consisting of a second plurality of sub-band artificial neural networks, each sub-band artificial neural network of the second plurality of sub-band artificial neural networks generating sub-band samples of a subset of artificial neural networks of the first plurality of artificial neural networks.

[0040] A particular advantage of this approach is that it allows for highly efficient sub-band processing, thereby allowing the required processing to be split across multiple (potentially smaller) artificial neural networks, which typically results in reduced complexity and / or an improved multi-channel audio signal.

[0041] According to an optional feature of the invention, the subband samples of the second subband samples of at least one artificial neural network of the second plurality of artificial neural networks include a plurality of subband samples of a plurality of processing time intervals of an artificial neural network of the first plurality of artificial neural networks.

[0042] This results in a particularly efficient implementation and / or improved performance.

[0043] According to an optional feature of the invention, the subband samples of the second subband samples of at least one artificial neural network of the second plurality of artificial neural networks comprise at least one subband sample of a subband of the subband representation of the downmix audio signal for which the at least one artificial neural network does not generate a subband sample of the subband representation of the auxiliary audio signal.

[0044] This results in a particularly efficient implementation and / or improved performance.

[0045] According to an optional feature of the invention, the first and second artificial neural networks are trained by a joint training process based on training data comprising a set of samples of a downmix audio signal generated by downmixing a training multi-channel audio signal and a target audio signal determined from a residual signal generated for the downmix audio signal, and using a cost function indicative of the difference between the auxiliary audio signal generated for the training multi-channel audio signal and the target audio signal.

[0046] This results in a particularly efficient implementation and / or improved performance.

[0047] According to an optional feature of the invention, the apparatus includes generating features for the set of features from an analytical analysis of the downmix audio signal.

[0048] This results in a particularly efficient implementation and / or improved performance.

[0049] According to another aspect of the present invention, there is provided a method for generating a multi-channel audio signal, the method comprising the steps of receiving a downmix audio signal of the multi-channel audio signal and upmix parametric data for upmixing the downmix audio signal, generating a set of features for the downmix audio signal with a first artificial neural network, the first artificial neural network having an input node for receiving a first sample of the downmix audio signal and an output node for providing the set of features, and a second artificial neural network having an input node for receiving a second sample of the downmix audio signal and an output node for providing a sample of an auxiliary audio signal for the downmix audio signal, the second artificial neural network further comprising a node for receiving a feature from the set of features, and generating the multi-channel audio signal from the downmix signal and the auxiliary audio signal depending on the upmix parametric data.

[0050] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. [Brief explanation of the drawings]

[0051] Embodiments of the invention will now be described, by way of example only, with reference to the drawings in which:

[0052] [Figure 1] FIG. 1 illustrates some elements of an example audio device according to some embodiments of the present invention. [Figure 2] Figure 2 shows an example of the structure of an artificial neural network. [Figure 3]FIG. 3 shows an example of a node in an artificial neural network. [Figure 4] FIG. 4 illustrates some elements of an example audio device according to some embodiments of the present invention. [Figure 5] FIG. 5 illustrates some elements of an example audio device according to some embodiments of the present invention. [Figure 6] FIG. 6 illustrates some elements of an example of an apparatus for training an artificial neural network of an audio device according to some embodiments of the present invention. [Figure 7] FIG. 7 illustrates some elements of a possible configuration of a processor for implementing elements of an audio device according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0053] FIG. 1 shows some elements of an audio device according to some embodiments of the present invention.

[0054] The audio device includes a receiver 101. This receiver 101 receives a data signal / bitstream comprising a downmix audio signal that is a downmix of a multi-channel audio signal. The following description focuses on the case where the multi-channel audio signal is a stereo signal and the downmix signal is a mono signal, but it will be understood that the techniques and principles described are equally applicable to multi-channel audio signals having more than two channels and to downmix signals having more than one channel (but fewer channels than the multi-channel audio signal).

[0055] Furthermore, the received data signal includes upmix parametric data for upmixing the downmix audio signal. Specifically, the upmix parametric data is a set of parameters that indicate relationships between signals of different audio channels of a multi-channel audio signal (specifically, a stereo signal) and / or between the downmix signal and an audio channel of the multi-channel audio signal. The upmix parameters typically indicate time differences, phase differences, level / intensity differences, and / or similarity measures (such as correlation). The upmix parameters are typically provided per time and per frequency (time-frequency tile). For example, new parameters are provided periodically for a set of subbands. In particular, the parameters include inter-channel phase difference (IPD), overall phase difference (OPD), inter-channel correlation (ICC), and channel phase difference (CPD) parameters, as known from parametric stereo coding (and higher channel coding).

[0056] Typically, the downmix audio signal is coded and the receiver 101 includes a decoder 103 for decoding the downmix audio signal, i.e., in this example, a mono signal. It will be understood that if the received downmix audio signal is not coded, the decoder 103 may not be required and may be considered to be an integral part of the receiver 101.

[0057] The receiver 101 is coupled to a generator 105 that generates a multi-channel audio signal from the downmix signal. The generator 105 generates the multi-channel audio signal from the downmix audio signal and the auxiliary audio signal depending on the parametric upmix data. The generator generates the output multi-channel audio signal by applying a 2×2 matrix multiplication to samples of the downmix audio signal and the auxiliary audio signal, particularly in the stereo case. The coefficients of the 2×2 matrix are determined from upmix parameters of the upmix parametric data, typically based on time and frequency bands. For other upmix operations, such as from a mono or stereo downmix signal to a five-channel multi-channel audio signal, the generator 105 can apply matrix multiplication with a matrix of appropriate dimensions.

[0058] It will be appreciated that many different techniques are known to those skilled in the art for generating such a multi-channel audio signal from a downmix audio signal and an auxiliary audio signal and for determining appropriate matrix coefficients from upmix parametric data, and any suitable technique may be used. In particular, various techniques for parametric stereo upmixing based on a downmix audio signal and an auxiliary audio signal are well known to those skilled in the art.

[0059] In conventional systems, upmixing involves generating an auxiliary audio signal in the form of a decorrelated version of the mono audio signal. Generating a decorrelated signal and mixing it with the mono audio signal has been found to improve the quality of the upmix signal, and decoders have been developed to take advantage of this. The decorrelated signal is typically generated by a decorrelator in the form of an all-phase filter applied to the mono audio signal. However, while the use of such an all-pass filter tends to produce a multi-channel audio signal that is perceived as having improved quality, it is still not ideal, and some degradation in audio quality is often noticeable.

[0060] The audio device of FIG. 1 uses techniques that have been found to tend to improve perceived audio quality in many scenarios and for many different audio signals.

[0061] In this approach, rather than generating a decorrelated signal by simple filtering of the downmix / mono audio signal, an auxiliary audio signal is generated by a specific configuration of a trained artificial neural network, which is used by the generator 105 to generate a multi-channel audio signal based on the upmix parameters.

[0062] The apparatus of FIG. 1 includes, among other things, a first artificial neural network (107) that receives samples of the downmix audio signal. These samples are fed to input nodes of the first artificial neural network. Output nodes of the first artificial neural network (107) provide a set of features of the downmix audio signal. The set of features may include several values ​​(e.g., scalar values) that reflect properties or characteristics of the downmix audio signal. Typically, the first artificial neural network (107) includes a much larger number of input nodes than output nodes. Therefore, a relatively small number of features are generated from a relatively large number of samples. The features thus provide a highly compressed and reduced representation of some properties or characteristics of the downmix audio signal. For example, the first artificial neural network (107) may have 1028 input nodes and 16 output nodes, thereby providing a highly compressed set of values ​​that depend on the properties of the downmix audio signal.

[0063] The apparatus further comprises a second artificial neural network 109 coupled to the first neural network and to the decoder 103 / receiver 105. The second artificial neural network 109 comprises, inter alia, an input node for receiving a sample of the downmix audio signal. The second artificial neural network 109 further comprises a node for receiving contributions from features of the set of features generated by the first artificial neural network 107. Such nodes may be input nodes of an input layer that also includes a node for receiving a sample of the downmix audio signal, or one, several or all of the nodes for receiving the features may be nodes of a layer different from the input layer of the downmix audio signal. For example, some or all of the nodes for receiving contributions from features are part of a hidden layer or a processing layer of the second artificial neural network 109.

[0064] The second artificial neural network 109 has output nodes that provide samples of the auxiliary audio signal, which are then fed to the generator 105 where the upmix operation is completed.

[0065] Thus, in this approach, the upmix process is not based on applying a decorrelation filter to the downmix audio signal to generate a decorrelated signal. This decorrelated signal is then combined with the downmix audio signal to generate a multi-channel audio signal. Rather, a trained artificial neural network structure generates an auxiliary audio signal that specifically replaces the decorrelated signal used in conventional upmix decoders. Both the first artificial neural network 107 and the second artificial neural network 109 have input nodes that receive samples of the downmix audio signal. In addition, the output of the first artificial neural network 107 is fed into the second artificial neural network 109 and used to control and adapt the processing of the second artificial neural network 109. Therefore, even though the weights and coefficients are constant, the second artificial neural network 109 can be considered not simply a fixed filter or trained network, but an adaptive or variable operation that is adapted based on the results of the first artificial neural network 107.

[0066] As will be explained in more detail below, various techniques can be used to train the neural network, and in particular, the overall training seeks to ensure that the output of the audio device is a multi-channel audio signal that most closely corresponds to the original multi-channel audio signal. Thus, the configuration is trained to provide an auxiliary audio signal that most effectively results in an accurate reconstruction of the multi-channel audio signal. In contrast to conventional approaches, such a signal is not necessarily a de-correlated signal. Rather, the second artificial neural network 109 is trained to generate an auxiliary audio signal that is optimal for combining with the downmix audio signal to generate the multi-channel audio signal. Such a signal is typically not a de-correlated version of the downmix audio signal, but rather a partially correlated signal, and in fact is likely to be close to the actual residual signal resulting from the original downmixing of the multi-channel audio signal. Thus, the user of the trained artificial neural network can enable the decoder to inherently and automatically account for and compensate for effects that may be introduced on the encoder side.

[0067] Similarly, the first artificial neural network 107 and the features it generates are not features that specifically represent particular properties or characteristics of the signal that are important to humans. Rather, the first artificial neural network 107 is trained to automatically adapt to provide features that are particularly suitable for adapting the second artificial neural network 109 to provide improved output values ​​for accurate reproduction of multi-channel audio signals.

[0068] In this approach, the artificial neural network architecture is appropriately configured to generate a second, auxiliary "augmentation" signal that aids and enhances the multi-channel reconstruction. For example, in a stereo signal, the encoder may generate a downmix signal as c *(l+r), where l and r represent the left and right channel signals, respectively, and c represents a time- and frequency-dependent scaling factor. The corresponding second signal for ideal reconstruction is d * (lr), where d is again time- and frequency-dependent. These two signals are not necessarily perfectly decorrelated, and a substantial advantage of the described approach is that, as opposed to simply attempting to decorrelate the mono downmix, the artificial neural network configuration generates an auxiliary audio signal that tends to approach the ideal signal d*(lr). This typically provides a significantly improved reconstruction of the original multi-channel audio signal.

[0069] The generation of the auxiliary audio signal is further improved by adapting a second artificial neural network 109 that generates this signal based on the set of features generated by the first artificial neural network 107. This adaptation has been shown to significantly improve reconstruction compared to a scenario that does not include adaptation.

[0070] Furthermore, the particular configuration allows for a very efficient implementation that can be achieved with relatively low complexity and resource usage.

[0071] The artificial neural network used in the above functions is a network of nodes arranged in layers, each node holding a node value. Figure 2 shows an example of a portion of an artificial neural network.

[0072] The node value of a given node is calculated to include contributions from some, or often all, of the nodes in the previous layer of the artificial neural network. Specifically, the node value of a node is calculated as a weighted sum of the node values ​​of all nodes output from the previous layer. A bias is typically added, and the result is influenced by an activation function. The activation function typically provides nonlinearity to account for the essential part of each neuron. Such nonlinearity and activation function have an important effect on the learning and adaptation process of the neural network. Therefore, the node value is generated as a function of the node values ​​in the previous layer.

[0073] Specifically, the artificial neural network includes an input layer 201 that includes a plurality of nodes that receive input data values ​​for the artificial neural network. Thus, the node values ​​of the nodes in the input layer are typically directly input data values ​​to the artificial neural network, and therefore are not calculated from other node values.

[0074] An artificial neural network may further include zero, one, or more hidden or processing layers 203. For each such layer, node values ​​are typically generated as a function of the node values ​​of the nodes in the previous layer, and an activation function (such as a sigmoid, ReLU, or Tanh function) is applied, in particular after weighting connections and adding biases.

[0075] Specifically, as shown in Figure 3, each node, also called a neuron, receives input values ​​(from nodes in the previous layer) and calculates the node value as a function of these values. Often, this involves first generating a value as a linear combination of the input values, where each input value is represented by a weight.

number

[0076] An activation function is then applied to the resulting connections. For example, the node value 1 is determined as follows: l=f(k) where the function is, for example, the normalized linear unit (as described in Xavier Glorot, Antoine Bordes, Yoshua Bengio, "Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics", PMLR15:315-323, 2011) function: f(k) = ReLU(k) = max(0,k)

[0077] Other commonly used functions include the sigmoid function and the tanh function. In many embodiments, the node output or value is calculated using multiple functions. For example, both the ReLU function and the sigmoid function can be combined using an activation function as follows: f(k)=ReLU(k)+σ(k)

[0078] Such operations are performed by each node of the artificial neural network (usually except for the input node).

[0079] The artificial neural network further includes an output layer 205, which provides output from the artificial neural network. That is, the output data of the artificial neural network are the node values ​​of the output layer. For hidden or processing layers, the output node values ​​are generated by functions of the node values ​​of previous layers. However, in contrast to hidden / processing layers, where the node values ​​are typically inaccessible or not used further, the node values ​​of the output layer are accessible and provide the results of the operation of the artificial neural network.

[0080] Several different network structures and toolboxes for artificial neural networks have been developed, and in many embodiments, artificial neural networks are based on the adaptation and customization of such networks. One example of a network architecture suitable for the above applications is WaveNet by van den Oord et al., which is described in Oord, Aaron van den, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, "Wavenet: A generative model for raw audio." arXiv preprint arXiv:1609.03499 (2016).

[0081] WaveNet is an architecture used for synthesis of time-domain signals using extended causal convolutions, and has been successfully applied to audio signals. The following activation functions are commonly used in WaveNet:

number

number

[0082] In some cases, the artificial neural network is further configured to include additional contributions that allow the artificial neural network to be dynamically adapted or customized to specific desired properties or characteristics of the generated output. For example, a set of values ​​is provided to adapt the artificial neural network. These values ​​are included by providing contributions to some nodes of the artificial neural network. These nodes are particularly input nodes, but may also typically be nodes in hidden or processing layers. Such adaptation values ​​are, for example, weighted and added as contributions to the weighted sum / correlation value of a given node. For example, in WaveNet, such adaptation values ​​are included in an activation function. For example, the output of the activation function may be given as:

number

[0083] The above description relates to neural network techniques that are suitable for many embodiments and implementations. However, it will be appreciated that many other types and structures of neural networks can be used. Indeed, many different techniques for generating neural networks have been developed and are being developed, including neural networks using different and more complex structures and processes. The techniques are not limited to any particular neural network technique, and any suitable technique can be used without detracting from the invention.

[0084] In many embodiments, the audio device performs subband processing. Specifically, as shown in Figure 4, the device of Figure 1 is modified to include a filter bank that generates a frequency or subband representation of the downmix audio signal. The filter bank may be a quadrature mirror filter (QMF) bank or may be implemented, for example, by a fast Fourier transform (FFT), although it will be appreciated that many other filter banks and techniques for splitting an audio signal into multiple subband signals are known and can be used. The filter bank may specifically be a complex-valued pseudo-QMF bank, resulting in, for example, 32 or 64 complex-valued subband signals.

[0085] In many embodiments, filter bank 401 generates a set of subband signals for subbands with equal bandwidths. In other embodiments, filter bank 401 generates subband signals with subbands having different bandwidths. For example, high-frequency subbands have higher bandwidths than low-frequency subbands. Subbands can also be grouped together to form a high-bandwidth subband.

[0086] Typically, the subbands have bandwidths ranging from 10 Hz to 10,000 Hz.

[0087] In some such embodiments, the artificial neural network that generates the samples of the auxiliary audio signal receives subband samples, i.e., samples of the subband audio signal. In particular, the apparatus of Figure 4 includes a plurality of subband neural networks 109, each receiving subband samples for a subband generated by a filter bank from the downmix signal. Each of the subband artificial neural networks further includes a set of features and then generates subband samples of the multi-channel audio signal for that subband.

[0088] The subband samples from the subband artificial neural network are provided to generator 105, which then generates a reconstructed multi-channel audio signal. For example, in some embodiments where a subband representation of the multi-channel audio signal is desired (e.g., because subsequent processing is also subband-based), generator 105 simply outputs the subband samples, possibly according to a particular structure or format, from the subband artificial neural network. In many embodiments, generator 105 is capable of converting the subband representation of the reconstructed multi-channel audio signal into a time-domain representation. In particular, generator 105 includes a synthesis filterbank that performs the inverse operation of filterbank 401, thereby converting the subband representation into a time-domain representation of the multi-channel audio signal.

[0089] Specifically, the generator generates the frequency / subband domain representation of the multi-channel audio signal by processing the frequency or subband domain representation of the downmix audio signal and the frequency / subband domain representation of the auxiliary audio signal. Thus, the processing of the generator 105 is subband processing, e.g., matrix multiplication performed on each subband of the subband samples of the downmix audio signal and the auxiliary audio signal generated by the corresponding subband artificial neural network.

[0090] The resulting subband / frequency domain representation can be used directly or can be converted to a time domain representation, for example using an appropriate synthesis filter bank, in particular applying a separate synthesis filter to each channel.

[0091] Each of the plurality of subband neural networks may be referred to as an auxiliary subband (region) artificial neural network, or more simply as an auxiliary subband artificial neural network. The comments made above with respect to the second artificial neural network 109 also apply mutatis mutandis to the auxiliary subband artificial neural network. Indeed, the second artificial neural network 109 may be considered one of the auxiliary subband artificial neural networks.

[0092] Thus, in this configuration, each of the auxiliary subband artificial neural networks receives subband samples of its subband, and further, all of the auxiliary subband artificial neural networks receive a set of features from the first artificial neural network 107.

[0093] Each of the auxiliary subband artificial neural networks generates subband samples for a subset of the subbands of the frequency subband representation of the auxiliary audio signal, typically generating subband samples (only) for the subbands that receive input samples from the filter bank 401 .

[0094] In many embodiments, the device has an artificial neural network for each subband of the frequency subband representation of the auxiliary audio signal generated by the filter bank. Thus, in many embodiments, the output samples for each subband of the filter bank 401 are fed to the input nodes of one auxiliary subband artificial neural network, which then generates subband samples of the auxiliary audio signal for that subband. Thus, in many embodiments, the subband processing is completely separate for each subband.

[0095] However, it will be understood that in some embodiments, auxiliary subband artificial neural networks may be provided for zero, one, or only some subbands, while other subbands may use auxiliary subband artificial neural networks that receive samples from multiple subbands, or indeed some subbands may not have auxiliary subband artificial neural networks applied (e.g., traditional decorrelation may be used for some subbands, as is typical for higher subbands).

[0096] Thus, in this embodiment, the generation of the auxiliary audio signal is performed on a subband-by-subband basis, with a separate and individual artificial neural network being used for each subband, each trained to provide output samples for the subband to which it is provided input subband samples, but the subband artificial neural network is further adapted based on the features generated by the first artificial neural network 107.

[0097] Such an approach has been found to provide highly advantageous generation of auxiliary audio signals that enable very high-quality reconstruction of multi-channel audio signals. Furthermore, the complexity is significantly reduced, typically allowing for highly efficient operations with significantly reduced computational resource requirements. Subband artificial neural networks tend to be significantly smaller than a single, complete artificial neural network required to generate the entire signal. Because processing typically requires far fewer nodes, and in some cases even fewer layers, the number of operations and calculations required to implement the artificial neural network function is significantly reduced. While more artificial neural networks are required to cover all subbands, smaller artificial neural networks typically significantly reduce the total number of operations required, and therefore the overall computational resource requirements. Furthermore, in many scenarios, a more efficient training process is possible.

[0098] The subband structure therefore provides a computationally efficient approach to enable an artificial neural network to be implemented to assist in the decoding of audio data including downmix audio signals and upmix parametric data. The described system and approach allows for the reconstruction of high quality multi-channel audio signals, typically achieving significantly improved audio quality compared to conventional approaches. Furthermore, a computationally efficient decoding process can be achieved. The subband and artificial neural network based approach is also compatible with other processes that use subband processing.

[0099] In some embodiments, subband processing is more adaptive than strict subband by subband processing. For example, in some embodiments, each auxiliary subband artificial neural network may receive subband samples not only from the subband itself, but also, possibly, from one or more other subbands. For example, in some embodiments, the auxiliary subband artificial neural network of one subband also receives samples of the downmix audio signal from one or two neighboring / adjacent subbands. As another example, in some embodiments, one or more of the auxiliary subband artificial neural networks also receives input samples from one or more subbands that contain harmonics (or subharmonics) of the frequency of the subband. For example, a subband around a center frequency of 500 Hz may also receive frequencies from a subband around a center frequency of 1000 Hz. Such additional subbands, which have a specific relationship with the subbands of the auxiliary subband artificial neural network, provide additional information that allows for generating an improved auxiliary subband artificial neural network for some audio signals.

[0100] In some embodiments, all auxiliary subband artificial neural networks have the same properties and dimensions. In particular, in many embodiments, all auxiliary subband artificial neural networks have the same number of input and output nodes, and possibly the same internal structure. Such an approach can be used, for example, in embodiments where all subbands have the same bandwidth.

[0101] However, in some embodiments, the auxiliary sub-band artificial neural networks may include non-identical neural networks. In particular, in some embodiments, the number of input nodes of the auxiliary sub-band artificial neural networks may differ for at least two of the artificial neural networks. Thus, in some embodiments, the number of input samples included in the determination of an output sample may differ for different sub-band and auxiliary sub-band artificial neural networks.

[0102] In some embodiments, the number of samples / input nodes may be greater in some low-frequency subbands than in some high-frequency bands. In fact, the number of samples / input nodes decreases monotonically with increasing frequency. Thus, the low-frequency auxiliary subband artificial neural network is larger and takes into account more input samples than the high-frequency auxiliary subband artificial neural network. This approach can be combined with subbands having different bandwidths, for example, where the low-frequency subband has a higher bandwidth than the high-frequency band.

[0103] In many scenarios, such an approach can improve the trade-off between achievable audio quality and computational complexity and resource usage, allowing the system to more closely adapt to reflect the typical characteristics of audio, thereby enabling more efficient processing.

[0104] In the above approach, all auxiliary subband artificial neural networks are provided with the same set of features. In some embodiments, only some of the auxiliary subband artificial neural networks are provided with the same set of features. However, in many embodiments, using the same set of features improves efficiency and performance. This can often reduce the complexity and resource usage in generating the feature sets. Furthermore, in many scenarios, it can speed up computation and allow each auxiliary subband artificial neural network to consider all available information provided by the features, thus improving the adaptation of the auxiliary subband artificial neural network.

[0105] However, in some embodiments, different sets of features may be provided to different artificial neural networks. For example, the first artificial neural network 107 generates a set of feature data values, and different subsets of these are provided to different auxiliary subband artificial neural networks. In other embodiments, several feature sets may also be generated, for example, including several features that are provided manually or generated by analysis of the downmix audio signal. For example, harmonics or peaks in the downmix audio signal may be detected. Such data may be applied, for example, to only some of the auxiliary subband artificial neural networks. For example, detected peaks or harmonics may be presented only to the auxiliary subband artificial neural networks of the subbands in which they were detected.

[0106] Thus, in many embodiments, different sets of features are provided to different auxiliary subband artificial neural networks, often involving some features being the same and some features being different for the different sets of auxiliary subband artificial neural networks.

[0107] In the above approach, a single first artificial neural network 107 is used to generate features for a set of features, while multiple auxiliary sub-band artificial neural networks process the sub-band downmix audio signal.

[0108] In other embodiments, the generation of the feature sets is subband-based. In some embodiments, separate filter banks are applied to the downmix audio signal to provide a set of subband signals. The sets of subband signals are then fed to corresponding subband artificial neural networks, each of which generates a set of features. Such multiple artificial neural networks are also referred to as feature set artificial neural networks. The feature sets generated by the feature set artificial neural networks are then fed to auxiliary subband artificial neural networks. The subbands generated by such filter banks need not be the same as the subbands generated by the filter bank 401 that generates the subband samples for the auxiliary subband artificial neural networks. For example, each of the subbands for determining the feature sets may include a different number of subbands for the auxiliary subband artificial neural networks, and the feature sets determined by one feature set artificial neural network may be fed to the appropriate auxiliary subband artificial neural network.

[0109] However, in many embodiments, the subbands used to generate the features are the same as the subbands used for the auxiliary subband artificial neural networks. Figure 5 shows an example of such an approach. In this example, rather than a single artificial neural network 107 generating a set of features, the apparatus of Figure 1 has been modified to include multiple artificial neural networks, each generating a set of features.

[0110] The comments made above with respect to the first artificial neural network 107 also apply mutatis mutandis to the feature set artificial neural network, and indeed the first artificial neural network 107 can be considered a feature set artificial neural network.

[0111] In this approach, each feature set artificial neural network generates a set of features that are applied to only a subset of the auxiliary subband artificial neural networks. Indeed, in many embodiments, each feature set artificial neural network generates a set of features for one of the auxiliary subband artificial neural networks. In some embodiments, the apparatus includes an equal number of feature set artificial neural networks and auxiliary subband artificial neural networks. In particular, there is one feature set artificial neural network and one auxiliary subband artificial neural network for each subband of the filter bank 401. In other embodiments, there are different numbers of feature set artificial neural networks and auxiliary subband artificial neural networks. For example, one feature set artificial neural network receives input samples for a group of subbands and generates a set of features for that group of subbands. This set of features is then applied to the group of auxiliary subband artificial neural networks for those subbands.

[0112] Such a subband-based approach to generating feature sets provides improved results in many scenarios by allowing more accurate feature sets to be generated and used to adapt the auxiliary subband artificial neural networks. It also reduces complexity and / or resource usage in many scenarios. For example, in many embodiments, the number of values ​​in the feature sets can be reduced, thereby reducing the complexity of both the feature set artificial neural network and the auxiliary subband artificial neural network. Furthermore, the feature set artificial neural network typically has fewer inputs and is much smaller than a full-bandwidth artificial neural network.

[0113] In many such embodiments, each of the feature set artificial neural networks generates a set of features for a given subset of subbands, typically a single subband, based on subband samples for that subset of subbands. However, in addition, one or more of the feature set artificial neural networks may further include subbands from one or more other subbands. That is, the input to a feature set artificial neural network may have an input node that receives subband samples for a subband for which the feature set artificial neural network does not generate a set of features.

[0114] As a specific example, in many embodiments, each feature set artificial neural network may receive as input not only the subband that produces the feature set, but also subsamples from, for example, adjacent subbands.

[0115] Such techniques often produce an improved set of features, which leads to improved audio quality. In particular, it has been found that the set of features can better reflect the temporal resolution of the downmix audio signal / multi-channel audio signal when surrounding subbands are taken into account. It has been found that such techniques can enable a better representation of temporal peakedness in particular.

[0116] In some embodiments, one or more of the feature set artificial neural networks, or indeed the first artificial neural network 107 if only one such artificial neural network is included, further includes as input subband samples from outside the time interval during which the corresponding auxiliary subband artificial neural network generates samples of the auxiliary audio signal.

[0117] In particular, the audio device processing is performed frame-by-frame, processing a time interval / frame of the received downmix audio signal to generate output samples of the multi-channel audio signal for that time interval / frame. Thus, for each frame, the filter bank 401 generates subband samples, which are fed to a feature set artificial neural network that generates a set of features for the time interval and a feature set artificial neural network that generates subband samples of a subband representation of the multi-channel audio signal based on the subbands and the set of features.

[0118] Thus, in particular, each feature set artificial neural network performs each operation in a block manner, where each operation generates a set of output samples from a set of input samples corresponding to a time interval of the downmix audio signal / multi-channel audio signal for which an output sample of the multi-channel audio signal is generated.

[0119] In some embodiments, one or more of the artificial neural networks receives, in addition to the appropriate subband samples generated for the current time interval, subband samples from other time intervals, typically from one or more adjacent time intervals. For example, in some embodiments, one or more of the feature set artificial neural networks also includes subband samples from the previous and next time intervals.

[0120] In many embodiments, such techniques can produce an improved set of features that improve audio quality.

[0121] Artificial neural networks are adapted to specific purposes through a training process used to adapt / tune / change the weights and other parameters (e.g., biases) of the artificial neural network. It will be appreciated that many different training processes and algorithms are known for training artificial neural networks. Typically, training is based on a large training set, where a large number of input data examples are provided to the network. Furthermore, the output of the artificial neural network is usually compared (directly or indirectly) to expected or ideal results. A cost function is generated to reflect the desired outcome of the training process. In a typical scenario known as supervised learning, the cost function often represents the distance between the prediction for specific input data and the ground truth. The weights can be modified based on the cost function, and by repeating the process with modified weights, the artificial neural network can be adapted toward a state where the cost function is minimized.

[0122] More specifically, during the training step, neural networks have two distinct information flows: from input to output (forward pass) and from output to input (backward pass). In the forward pass, data is processed by the neural network as described above, while in the backward pass, weights are updated to minimize a cost function. Typically, such backward propagation follows the gradient direction of the cost function landscape. In other words, for a batch of input data, by comparing the predicted output with the ground truth, the direction in which the cost function will be minimized and backward propagation can be estimated by appropriately updating the weights. Other known techniques for training artificial neural networks include, for example, the Levenberg-Marquardt algorithm, the conjugate gradient method, and Newton's method.

[0123] In this case, specifically, the training includes a training set containing a potentially large number of multi-channel audio signals or corresponding downmix audio signals. The training set includes audio signals representing a large number of different audio sources, including recordings of videos, movies, telecommunications, etc. In some embodiments, the training data may include non-audio data, such as when training is performed in combination with training data from other sources, such as text data.

[0124] In some embodiments, the training data is a multi-channel audio signal in a time segment corresponding to the processing time interval of the artificial neural network being trained. For example, the number of samples in the training multi-channel audio signal corresponds to the number of samples corresponding to the input nodes of the artificial neural network being trained. Thus, each training example corresponds to one operation of the artificial neural network being trained. However, typically, to speed up the training process, batches of training samples are considered at each step. Furthermore, many upgrades to gradient descent are possible to speed up convergence or avoid local minima in the cost function landscape.

[0125] For each training multi-channel audio signal, the training processor performs a downmix operation to generate a downmix audio signal and corresponding upmix parametric data, such that during normal operation, the encoding process applied to the multi-channel audio signals is also applied to the training multi-channel audio signals, thereby generating the downmix and upmix parametric data.

[0126] Furthermore, in some embodiments, the training processor generates a residual signal that reflects differences between the downmix audio signal and the multi-channel audio signal, or more typically, represents parts of the multi-channel audio signal that are not adequately represented in the downmix audio signal. For example, in many embodiments, the training processor generates a downmix signal and, in addition thereto, generates a residual signal. The residual signal, when used in upmixing based on the upmix parametric data, can reconstruct a (more) accurate multi-channel audio signal.

[0127] In particular, for stereo multi-channel audio signals, the training processor can use a parametric stereo scheme (e.g., according to an appropriate standardized technique). Such encoding applies frequency- and time-dependent matrix operations, e.g., rotation operations, to the input stereo signal to generate a downmix signal and a residual signal. For example, a 2×2 matrix / complex-value multiplication is typically applied to the input stereo signal to substantially align, e.g., one of the rotated channel signals so that it has the maximum signal value. This channel is used as a mono signal, and the rotation is typically performed on a frame basis. The rotation value is stored as part of the upmix parametric data (or a parameter enabling this determination may be included in the upmix parametric data). The synthesizer then performs an inverse rotation to reconstruct the stereo signal. The rotation of the stereo signal results in another stereo signal in which one channel is properly aligned to maximum intensity. This other channel is typically discarded in the parametric stereo encoder to reduce the data rate. In conventional PS decoding, a decorrelated signal is typically generated in the decoder and used in the upmixing process. In current training techniques, this second signal is used as the residual signal for downmixing, as it may represent information discarded at the encoder, and therefore represents the ideal signal to be reconstructed at the decoder as part of the upmixing process.

[0128] Thus, in some embodiments, the training processor generates a training downmix signal and / or a training residual signal from the training multi-channel audio signal. The training downmix signal is fed to the construction of the first artificial neural network 107 and the second artificial neural network 109, or equivalently to the construction of the feature set artificial neural network and the auxiliary subband artificial neural network. That is, samples of the training downmix audio signal are fed to the neural networks using the same processing as is applied to the downmix audio signal by the audio device during normal operation (including, for example, subband filtering, etc.).

[0129] Next, an output from the neural network operation is determined, and a cost function is applied to determine a cost value for each training downmix audio signal and / or a combination set of training downmix audio signals (e.g., determine an average cost value of the training set). The cost function may include various components.

[0130] Typically, the cost function includes at least one component that reflects how close the generated signal is to the reference signal, i.e., the so-called reconstruction error. In some embodiments, the cost function includes at least one component that reflects how close the generated signal is to the reference signal from a perceptual point of view.

[0131] For example, in some embodiments, the auxiliary audio signal generated by the second artificial neural network 109 for a given training downmix audio signal / multi-channel audio signal is compared with the residual signal of that training downmix audio signal / multi-channel audio signal. A cost function contribution / combination is generated that reflects the difference between the generated auxiliary audio signal and the reference residual signal. This process is performed for all training sets to generate an overall cost function.

[0132] One such example is shown in Figure 6. In this example, a downmixer 601 receives a training multi-channel audio signal, which in this example is a stereo signal. For the given training signal, the downmixer 601 performs a downmixing operation to generate a training downmix signal, which in this example is a training mono audio signal and a residual signal. The downmix audio signal is provided to a preprocessor 603, which generates downmix audio signal samples for input to the artificial neural network. In particular, the preprocessor 603 performs the same operations as those performed in a synthesis device to generate samples for the artificial neural network; in fact, the same functions are typically used. That is, the decoder function used to generate input samples for the artificial neural network is also used in the training process. Furthermore, the preprocessor 603 includes functions corresponding to the encoding / decoding processing of the training downmix audio signal, including, for example, quantization.

[0133] The output of the pre-processor 603 is provided to a first artificial neural network 107 and a second artificial neural network 109, similar to the audio device of Figure 1. The first artificial neural network 107 and the second artificial neural network 109 are coupled to each other such that the output set of features generated by the first artificial neural network 107 is provided to the second artificial neural network 109. The output of the second artificial neural network 109 therefore corresponds to samples of the auxiliary audio signal generated for the training downmix audio signal.

[0134] Thus, for a given training multi-channel audio signal, the training system of FIG. 6 performs the same operations as the audio device of FIG. 1, thereby generating the auxiliary audio signal that the audio device would generate if the artificial neural network had the same data / configuration (coefficients, biases, etc.).

[0135] The output of the second artificial neural network 109 is provided to a comparator 605, which compares the generated auxiliary audio signal with the residual signal generated by the downmixer 601. As mentioned above, the residual signal allows, in principle, a substantially perfect reconstruction of a multi-channel audio signal and can therefore be considered an approximation of an ideal auxiliary audio signal. Therefore, the comparison between the generated auxiliary audio signal and the residual signal indicates how advantageous the generated auxiliary audio signal is. Therefore, a cost value is determined based on the comparison; specifically, the greater the difference, the higher the cost value.

[0136] It will be appreciated that many different techniques can be used to determine a cost value that reflects the difference between the signals. For example, a correlation can be performed with a cost value having a monotonically decreasing value for increasing correlation values. As another example, two signals can be subtracted from each other and a power measure of the difference signal can be used as the cost value. It will be appreciated that many other techniques are available and can be used.

[0137] Thus, in this example, the cost function generates a cost value that reflects how closely the generated auxiliary audio signal matches the corresponding residual signal of the training multi-channel audio signal.

[0138] Based on the cost values, the training processor 607 adapts the weights of the artificial neural networks. For example, a backpropagation technique is used. In particular, the training processor 607 adjusts the weights of both the first artificial neural network 107 and the second artificial neural network 109 based on the cost values. For example, given the derivative of the weight with respect to the cost function (representing the slope), the weight value is changed to progress in the direction of the slope. A simple / minimal account can refer to training a perceptron (single neuron) in the case of a backward pass with a single data input.

[0139] This process is repeated until the artificial neural network is considered trained. For example, training is done for a predetermined number of iterations. As another example, training continues until the weight changes are less than a predetermined amount. Also, very commonly, validation stopping is implemented, where the network is tested against a validation metric and stopped when it reaches an expected result.

[0140] As a specific example, a stereo signal is fed into a conventional PS downmix module, which generates both a downmix and an ideal residual signal, i.e., a residual signal that allows for (near) perfect reconstruction of the waveform at the decoder side. Using the mono and residual signal pair, a first artificial neural network 107 generates a set of features that describe the audio signal (frame). This set of features is fed, along with the mono audio signal, to a second artificial neural network 109, which generates an auxiliary audio signal. When the artificial neural network is deep enough, this configuration is trained using a cost function, such as RMSE (root mean square error of the auxiliary audio signal relative to the residual signal), resulting in an artificial neural network configuration that learns what the auxiliary audio signal should be for a given mono audio signal.

[0141] In this configuration, the first artificial neural network 107 and the second artificial neural network 109 are suitably trained together using the same downmix audio signal and cost function, and the weights of both the first artificial neural network 107 and the second artificial neural network 109 are updated based on the same training data and the same cost function and downmix audio signal.

[0142] In many embodiments, the residual signal is used directly when comparing with the generated auxiliary audio signal. However, more commonly, a target audio signal generated from the residual audio signal is generated and used in comparison with the generated auxiliary audio signal. The target audio signal is generated by applying a function or signal processing application to the residual signal. For example, scaling / level setting of the residual audio signal can be applied to generate the target audio signal. As another example, a filter operation can be applied to the residual audio signal to generate the target audio signal. As another example, scaling can be applied to the residual signal to minimize the difference and / or maximize the correlation with the generated auxiliary audio signal.

[0143] In some embodiments, alternatively or additionally, the cost function reflects the difference between the training multi-channel audio signal and a multi-channel audio signal generated by upmixing the downmix audio signal and the generated auxiliary audio signal. In such an example, the downmixer 601 also generates upmix parametric data used in the upmixing. Thus, in some embodiments, rather than simply training an artificial neural network to generate an auxiliary audio signal that matches the residual signal, the training includes reconstructing the multi-channel audio signal based on the generated auxiliary audio signal. In particular, the generation of the output multi-channel audio signal is performed using the same operations as those performed in the encoder. This output multi-channel audio signal is then compared with the input multi-channel audio signal.

[0144] In some embodiments, the cost function also considers other parameters. For example, in some embodiments, the cost function also considers the degree of correlation between the generated auxiliary audio signal and the downmix signal. In particular, the cost function indicates a lower cost the more the downmix audio signal and the auxiliary audio signal are decorrelated. Increased decorrelation may indicate that less information is common to both of the two signals.

[0145] In embodiments where subband operations are performed, the described techniques are performed for each subband. Specifically, a residual audio signal is generated for each subband and compared with the auxiliary audio signal subband samples generated for that subband. Thus, a cost function is evaluated for each subband to determine a cost value for training the coefficients of that subband's feature set artificial neural network and auxiliary subband artificial neural network.

[0146] In many embodiments, the feature set artificial neural network, and / or indeed the first artificial neural network 107, generates a set of features that matches the time and frequency resolution of the auxiliary subband artificial neural network, including, in particular, the second artificial neural network 109 if only one artificial neural network is used. However, in other embodiments, the time and / or frequency resolution of the generated feature set may be different, typically a lower resolution for the feature set. In such situations, the ability to change the resolution may be included. In particular, an interpolator may be incorporated to generate a set of interpolated parameter values ​​from the generated parameter values.

[0147] In some embodiments, one or more of the sets of features includes, in addition to features generated by the feature set artificial neural network, one or more values ​​generated in other ways.

[0148] Indeed, in some embodiments, features reflecting user input may be included. In particular, the devices of FIGS. 1, 4, and 5 include a user input and a user input processor that generates features in response. This approach allows a user to control or adapt the operation of the generation of a multi-channel audio signal. For example, it allows the user to select a specific audio mode that can affect the reconstructed multi-channel audio signal by generating an auxiliary audio signal. The user input relates to a perceptual audio quality manually created by an expert. This additional term may reduce the quantitative reconstruction of the audio, expressed for example as RMSE, but may improve the human perception of the audio reconstruction.

[0149] Training of the auxiliary subband artificial neural network is performed, for example, depending on this feature set. For example, the features may have a discrete number of possible values. Each possible value corresponds to one user setting. Each of these settings corresponds to a desired variation of the multi-channel audio signal. For example, one user input may increase low and high frequencies, another may be filtered to provide only mid-range sounds, and a third setting may have a large amount of reverberation or echo. During training, a reference multi-channel audio signal used for comparison in the cost function is processed according to the specific preferences of the different settings. The network is then trained for all possible settings of the user mode features, but the reference multi-channel audio signal is selected as the one processed for that particular user mode feature.

[0150] As another example, features are generated in response to different modalities, for example, using face detection to identify users, and features are set to reflect the current user, thereby automatically adapting to individual users.

[0151] In many embodiments, one or more of the features are generated by analysis of the downmix audio signal.

[0152] Many different algorithms and procedures are known for analysing audio signals to extract or determine properties of the audio signals, and it will be appreciated that any suitable (analysis) techniques, algorithms and properties may be used.

[0153] In particular, the audio device may be implemented in one or more suitably programmed processors. In particular, the artificial neural network may be implemented in one or more such suitably programmed processors. The various functional blocks, in particular the artificial neural network, may be implemented in separate processors or, for example, in the same processor. An example of a suitable processor is shown below:

[0154] 7 is a block diagram illustrating an exemplary processor 700 according to an embodiment of the present disclosure. Processor 700 may be used to implement one or more processors that implement the aforementioned devices or elements thereof, including, inter alia, one or more artificial neural networks. Processor 700 may be any suitable processor type, including, but not limited to, a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable array (FPGA) (wherein the FPGA is programmed to form a processor), a graphics processing unit (GPU), an application specific integrated circuit (ASIC) (wherein the ASIC is designed to form a processor), or a combination thereof.

[0155] Processor 700 includes one or more cores 702. Core 702 includes one or more arithmetic logic units (ALUs) 704. In some embodiments, core 702 may include a floating point logic unit (FPLU) 706 and / or a digital signal processing unit (DSPU) 708 in addition to or instead of the ALUs 704.

[0156] The processor 700 includes one or more registers 312 communicatively coupled to the core 702. The registers 712 are implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 712 are implemented using static memory. The registers provide data, instructions, and addresses to the core 702.

[0157] In some embodiments, processor 700 includes one or more levels of cache memory 710 communicatively coupled to cores 702. Cache memory 710 provides computer-readable instructions to cores 702 for execution. Cache memory 710 provides data for processing by cores 702. In some embodiments, computer-readable instructions are provided to cache memory 710 by local memory (e.g., local memory attached to external bus 716). Cache memory 710 is implemented with any suitable cache memory type, such as, for example, metal-oxide-semiconductor (MOS) memory, such as, for example, static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology.

[0158] Processor 700 includes a controller 714. Controller 714 controls input to processor 700 from and / or output from other processors and / or components included in the system. Controller 714 controls data paths within ALU 704, FPLU 706, and / or DSPU 708. Controller 714 may be implemented as one or more state machines, data paths, and / or dedicated control logic. Gates in controller 714 may be implemented as stand-alone gates, FPGAs, ASICs, or any other suitable technology.

[0159] Registers 712 and cache 710 communicate with controller 714 and core 702 via internal connections 720A, 720B, 720C, and 720D. The internal connections may be implemented as buses, multiplexers, crossbar switches, and / or any other suitable connection technology.

[0160] Input and output for processor 700 is provided via bus 716. Bus 716 includes one or more conductive lines. Bus 716 is communicatively coupled to one or more components of processor 700, such as controller 714, cache 710, and / or registers 712. Bus 716 is coupled to one or more components of the system.

[0161] Bus 716 is coupled to one or more external memories. The external memories include read-only memory (ROM) 732. ROM 732 may be masked ROM, electronically programmable read-only memory (EPROM), or any other suitable technology. The external memories include random access memory (RAM) 733. RAM 733 may be static RAM, battery-backed static RAM, dynamic RAM (DRAM), or any other suitable technology. The external memories include electrically erasable programmable read-only memory (EEPROM) 735. The external memories include flash memory 734. The external memories include magnetic storage devices such as disks 736. In some embodiments, external memories may be included in the system.

[0162] The invention can be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention may optionally be implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the invention may be physically, functionally, and logically implemented in any suitable way. Indeed, functionality may be implemented in a single unit, in several units or as part of other functional units. Thus, the invention may be implemented in a single unit or may be physically and functionally distributed between different units, circuits, and processors.

[0163] While the present invention has been described in connection with some embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Moreover, while certain features may appear to be described in connection with particular embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.

[0164] Furthermore, although individually listed, a plurality of means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Moreover, although individual features may be included in different claims, these features may also be advantageously combined, and the inclusion of features in various claims does not imply that such combinations are not feasible and / or advantageous. Furthermore, the inclusion of a feature in one category of claims does not imply a limitation of this category, but rather indicates that the feature may be applied to other claim categories as well, where appropriate. Furthermore, the order of features in the claims does not imply a particular order in which the features must function, and in particular the order of individual steps in method claims does not imply that the steps must be performed in this order. Rather, steps may be performed in any suitable order. Furthermore, a reference to the singular does not exclude a reference to the plural. Thus, references to "first," "second," etc. do not exclude the plural. Reference signs in the claims are provided solely as examples for clarity, and these examples should not be construed as limiting the scope of the claims in any way.

Claims

1. A device for generating multi-channel audio signals, A receiver that receives a downmix audio signal of the multi-channel audio signal and upmix parametric data for upmixing the downmix audio signal, A first artificial neural network for generating a set of feature quantities of the downmixed audio signal, the first artificial neural network having an input node for receiving a first sample of the downmixed audio signal and an output node for providing the set of feature quantities, A second artificial neural network having an input node for receiving a second sample of the downmix audio signal and an output node for providing a sample of an auxiliary audio signal of the downmix audio signal, further comprising a node for receiving features from the set of features, A generator that generates the multi-channel audio signal from the downmix signal and the auxiliary audio signal, depending on the upmix parametric data, A device including a device.

2. The apparatus according to claim 1, comprising a first filter bank for generating frequency subband representations of the downmix audio signal, wherein at least some of the second samples of the downmix audio signal are subband samples of the frequency subband representation.

3. The apparatus according to claim 2, wherein the second artificial neural network is an artificial neural network comprising a plurality of first subband artificial neural networks, and each subband artificial neural network of the first plurality of subband artificial neural networks generates subband samples of a subset of subbands of the frequency subband representation of the auxiliary audio signal.

4. The apparatus according to claim 3, wherein the plurality of subband artificial neural networks include artificial neural networks for each subband of the frequency subband representation of the auxiliary audio signal.

5. The apparatus according to claim 3, wherein the generator generates a frequency subband representation of the multichannel audio signal by applying subband matrix operations to the frequency subband representation of the auxiliary audio signal and the frequency subband representation of the downmix audio signal, and converts the frequency subband representation of the multichannel audio signal into a time-domain representation of the multichannel audio signal.

6. The apparatus according to claim 3, wherein the set of features generated by the subband artificial neural network among the first plurality of subband artificial neural networks is common to the plurality of subbands of the frequency subband representation of the downmix audio signal.

7. The apparatus according to claim 3, wherein the number of input nodes of the artificial neural network among the first plurality of subband artificial neural networks decreases monotonically as the frequency increases.

8. The apparatus according to any one of claims 1 to 7, comprising a second filter bank for generating frequency subband representations of the downmix audio signal, wherein at least some of the first samples of the downmix audio signal are subband samples of the frequency subband representation.

9. The apparatus according to claim 8, which is dependent on claim 3, wherein the first artificial neural network is an artificial neural network comprising a second plurality of subband artificial neural networks, and each subband artificial neural network of the second plurality of subband artificial neural networks generates subband samples of a subset of artificial neural networks among the first plurality of artificial neural networks.

10. The apparatus according to claim 9, dependent on claim 3, wherein the subband samples of the second subband samples of at least one artificial neural network among the second plurality of artificial neural networks include a plurality of subband samples of a plurality of processing time intervals of the artificial neural network among the first plurality of artificial neural networks.

11. The apparatus according to claim 9 or 10, wherein the subband sample of the second subband sample of at least one of the second plurality of artificial neural networks includes at least one subband sample of the subband of the subband representation of the downmix audio signal, for which the at least one artificial neural network does not generate a subband sample of the subband representation of the auxiliary audio signal.

12. The apparatus according to claim 1, wherein the first artificial neural network and the second artificial neural network are trained by a joint training process based on training data including a set of samples of a downmix audio signal generated by downmixing a training multichannel audio signal and a target audio signal determined from residual signals generated for the downmix audio signal, and using a cost function that shows the difference between an auxiliary audio signal generated for the training multichannel audio signal and the target audio signal.

13. The apparatus according to claim 1, further comprising generating feature quantities for the set of feature quantities from analytical analysis of the downmix audio signal.

14. A method for generating a multi-channel audio signal, The steps include receiving a downmix audio signal of the multi-channel audio signal and upmix parametric data for upmixing the downmix audio signal, A first artificial neural network generates a set of features of the downmixed audio signal, wherein the first artificial neural network has an input node for receiving a first sample of the downmixed audio signal and an output node for providing the set of features. A second artificial neural network having an input node for receiving a second sample of the downmix audio signal and an output node for providing a sample of an auxiliary audio signal of the downmix audio signal, wherein the second artificial neural network further includes a node for receiving features from the set of features, The steps include generating the multi-channel audio signal from the downmix signal and the auxiliary audio signal, depending on the upmix parametric data, Methods that include...

15. A computer program comprising computer program code means, wherein the computer program code means is adapted to perform all the steps of the method according to claim 14 when the computer program is executed on a computer.