A system for automatic multitrack mixing.
A hybrid deep learning system for automatic multi-track mixing addresses data scarcity and training complexity by using a first network to learn signal processing algorithms and a second network to control sub-modules, facilitating efficient and high-quality audio mixing.
Patent Information
- Application Number
- JP2022578976
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-22
- Filing Date
- 2021-06-16
- Publication Date
- 2025-10-06
- Estimated Expiration
- 2041-06-16
AI Technical Summary
The lack of multitrack data and the complexity of training deep learning models for automatic multi-track mixing pose challenges in achieving efficient and effective audio mixing, particularly due to the diversity in mixing conventions and the need for supervised learning of a one-to-many mapping in mix space.
A hybrid deep learning system is proposed, comprising a first network that learns general signal processing algorithms of a mixing console and a second network that controls these sub-modules to generate a mix, allowing for weight-sharing and efficient training with a hybrid approach.
The system enables flexible and efficient automatic multi-track mixing, capable of generating high-quality stereo mixes with reduced data requirements and simplified training complexity.
Smart Images

Figure 0007749603000001 
Figure 0007749603000002 
Figure 0007749603000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to the field of audio mixing. In particular, the present disclosure relates to techniques for automatic multi-track mixing in the waveform domain using machine learning models or systems, and to a framework for training such machine learning models or systems. [Background technology]
[0002] Generally speaking, the path from the seed of a musical idea to a final recorded work involves many different steps that are often not obvious from the perspective of the music listener. This process typically requires the collaboration of many different individuals, each with their own unique roles and skills, such as songwriters, musicians, producers, and recording, mixing, and mastering engineers. One critical step within this process is the task of transforming the individual recorded elements into the final mixture, a task performed by the mixing engineer and often an integral part of the creative process in modern recording.
[0003] This task of transforming a collection of audio signals into a cohesive mixture requires a deep understanding of disparate technical and creative processes. To perform this task effectively, an audio engineer's professional training includes developing an ability to recognize how to utilize a set of signal processing tools to achieve a set of desired technical and creative goals. This reality has led to many drivers for the development of intelligent music production (IMP) tools, tools intended to assist in parts of this complex process.
[0004] In recent years, deep learning has demonstrated impressive results in many audio tasks previously considered extremely challenging. While these recent successes seem promising for our goal of advancing IMP systems, many challenges hinder our ability to design deep models for multitrack mixing tasks. Foremost among these is the limited availability of multitrack mix data. In addition to the lack of parallel data, there is also the challenge of building models that can adapt to the diversity present in real-world multitrack projects. A final, often overlooked challenge stems from the nature of the mixing task. While professionally produced mixes have conventions, it is clear that different but equally acceptable mixes may exist in disparate regions of the so-called mix space, or parameter space, of a mixing console. Even ignoring the task of attempting to model this complex one-to-many mapping, the inherent challenges of training models to regress a ground truth mix in a supervised manner remain. Summary of the Invention
[0005] Thus, there is a need for methods and systems for performing (automatic) multi-track audio mixing, and possibly methods for training systems for (automatic) multi-track mixing that can achieve improved performance (e.g., in terms of error rate, consistency, etc.) and / or efficiency, while at the same time allowing for good generalization to new audio (e.g., recordings) and / or listeners.
[0006] In view of the above, the present disclosure generally provides a deep learning based system for automatic multi-track mixing based on multiple input audio tracks, a method for operating a deep learning based system for automatic multi-track mixing based on multiple input audio tracks, and a method for training a deep learning based system for automatic multi-track mixing based on multiple input audio tracks, as well as corresponding (computer) programs, computer readable storage media, and devices, comprising the features of each independent claim. The dependent claims relate to preferred embodiments.
[0007] According to one aspect of the present disclosure, a deep learning-based system for automatic multi-track mixing based on multiple input audio tracks is provided. The input audio tracks may be, for example, pre-recorded or provided in real time. The input audio tracks may also undergo appropriate (pre-)processing, if necessary. The system may include one or more instances of a first deep learning-based network. In particular, the first network may be configured to generate parameters used for automatic multi-track mixing based on the input audio tracks. The parameters may include, but are not limited to, control parameters, panning parameters, or any other suitable parameters suitable for the audio mixing process. The system may further include one or more instances of a second deep learning-based network. The first network instance and / or the second network instance may be configured in a weight-sharing manner. That is, when multiple first and / or second deep learning-based networks are used, all of the first and / or second networks may be configured with the same weights (e.g., weight vectors). In particular, the second network is configured to apply signal processing and at least a mixing gain to the input audio tracks based on the parameters to generate an output mix of audio tracks. In this sense, in some cases, the first network may be referred to as a controller network (providing parameters such as control parameters), while the second network may be referred to as a transformation network (performing audio transformation or processing). Signal processing may refer to applying appropriate audio effects to the audio tracks. Signal processing (or audio signal effects) may include, but are not limited to, gain, panning, equalization, dynamic range compression, reverberation, etc., as will be understood by those skilled in the art. In the context of multi-track audio mixing, (mixing) gain may in some cases be considered a simple and basic transformation.Whereas in other cases, mixing gains and panning gains may be applied, but it should be understood that any other suitable audio processing implementation may be performed depending on different scenarios.
[0008] Broadly speaking, the proposed system architecture, configured as described above, aims to learn the mixing practices of audio engineers by directly observing the audio transformation from original instrument recordings to the final mix. Furthermore, to address the lack of multitrack data and the large amount of data typically required to train deep learning models, the proposed system architecture presents a hybrid approach. Generally speaking, the hybrid approach involves a first model (second network) that learns the general signal processing algorithms of a mixing console, followed by a second model (first network) that learns to control these channel-like submodules to generate the (final) mix. As such, a system for automatic multitrack mixing can be trained in a flexible and efficient manner, allowing the system to be utilized / operated to perform multitrack mixing on input audio tracks with reasonable mixing quality.
[0009] In some examples, the output mix may be a stereo mix. That is, the output mix may have one mix for the left channel and another mix for the right channel. In particular, in the case of stereo mixing, the second network may be configured to apply stereo mixing gains to the input audio tracks, i.e., one mixing gain for the left channel and another mixing gain for the right channel. Alternatively, the second network may also be configured to apply one mixing gain and one panning parameter (e.g., between 0 and 1) to generate the output stereo mix, which may simply mean that the parameter for the other channel simply corresponds to 1 minus the panning parameter.
[0010] In some examples, the first network and the second network may be trained separately. In particular, the second network may be trained first (i.e., before the first network is trained), so that the first network can be trained based on the pre-trained second network. Such a hybrid configuration in which the first network and the second network are trained separately can significantly simplify the complexity of implementing and training the entire mixing system.
[0011] In some examples, the number of instances of the first network and / or the number of instances of the second network may be determined in response to (or based on) the number of input audio tracks. For example, in a possible implementation, the number of instances of the second network may be equal to the number of input audio tracks. On the other hand, depending on various implementations, one or more instances of the first network may also be provided in response to the number of input audio tracks and / or the number of instances of the second network.
[0012] In some examples, the first network may have a first stage and a second stage. In particular, generating parameters by the first network may include mapping each of the input audio tracks to a respective feature space representation by the first stage, and generating parameters used by the second network based on the feature space representations by the second stage. The feature space representations may be, for example, latent space representations. In this sense, the first stage is sometimes referred to as an encoding stage (or in some cases simply as an encoder).
[0013] In some examples, generating parameters for use by the second network by the second stage may include generating a combined representation based on the feature space representation of the input audio track, and generating parameters for use by the second network based on the combined representation (in addition to or as an alternative to the feature space representation). The combined representation may be a concatenated representation of the feature space representation of the input audio track, which may be generated by any suitable means.
[0014] In some examples, generating the combined representation may involve (possibly among other things) an averaging process on the feature space representations of the input audio tracks.
[0015] In some examples, the first network may be trained based on at least one loss function that indicates the difference between a given mix of audio tracks and their respective predictions.
[0016] In some examples, the first network may be trained by using any suitable means. Training may refer to determining parameters for a deep learning model (e.g., a neural network) used to implement the system. Furthermore, training may refer to iterative training. In a possible implementation, training may include obtaining at least one first training set as input. In particular, the first training set may include multiple subsets of (e.g., pre-recorded) audio tracks, and for each subset, one or more predetermined mixes of the audio tracks in that subset. The predetermined mixes of the audio tracks in the subset may be provided (e.g., mixed), for example, by an audio mixing engineer or by any suitable means, such that the predetermined mixes of the audio tracks can represent appropriate and acceptable audio mixes that can serve as a suitable basis (target) for training. Furthermore, training may include inputting the first training set to the first network and iteratively training the first network to predict the mix of each of the subset's audio tracks in the training set. In particular, the training may be based on at least one first loss function that indicates the difference between a given mix of audio tracks and their respective predictions.
[0017] In some examples, the predicted mix of audio tracks may be a stereo mix. Accordingly, the first loss function may be a stereo loss function and may be configured to be invariant under left and right channel reassignments. In some possible cases, such invariance between stereo channels may be achieved by considering the sum of audio signals corresponding to the left and right channels, rather than considering them separately.
[0018] In some examples, training a first network to predict a mix of audio tracks may include, for each subset of audio tracks, generating, by the first network, a plurality of prediction parameters according to the subset of audio tracks, providing the prediction parameters to a second network, and generating, by the second network, a prediction of a mix of the subset of audio tracks based on the prediction parameters and the subset of audio tracks.
[0019] In some examples, the number of instances of the second network may be equal to the number of input audio tracks. In this case, the second network may be configured to perform signal processing on each input audio track based at least in part on the parameters to generate a respective processed output. In particular, the processed output may be a mono audio signal or may have a left channel and a right channel (i.e., a stereo audio signal). An output mix may be generated based on the processed (e.g., mono or stereo) output.
[0020] In some examples, the system may further include a routing component (e.g., a router). In particular, the routing component may be configured to generate multiple intermediate mixes (e.g., stereo mixes) based on the processed outputs, and then the output mix may be generated based on those intermediate mixes. In other words, the routing component may be considered to provide further appropriate audio signal processing before generating a final audio mix. In some possible implementations, the routing component may be configured to generate appropriate bus-level mixes, so that the final audio mix may be generated based on those bus-level mixes.
[0021] In some examples, the first network may be configured to further generate appropriate (e.g., control, routing, etc.) parameters for the routing component. For example, the first network may be configured to generate appropriate (operational) parameters for the routing component to perform bus level mixing.
[0022] In some examples, one or more instances of the second network may be referred to as a first set of one or more instances of the second network, and the system may further include a second set of one or more (weight-sharing) instances of the second network. In particular, the number of instances of the second set may be determined according to the number of intermediate mixes. For example, in a possible implementation, each of the intermediate mixes (e.g., bus-level mixes) generated by the routing component may be processed / manipulated by a stereo link pair of the second network, and for each of the intermediate mixes as input, a dual stereo output may be generated by the stereo link pair of the second network. Specifically, a stereo link connection may refer to a configuration of a pair of (two) second networks, each with its own input, where both second networks use the same parameters to apply the same signal processing to each input and each generate a stereo output. Dual stereo may refer to a configuration of the system where the input is a stereo signal and the output generates separate signals for left and right outputs (channels), which are later summed to generate an overall left and right stereo signal.
[0023] In some examples, the first network may be further configured to generate (eg, bus control) parameters for a second set of instances of the second network.
[0024] In some examples, the system may be configured to further generate a left mastering mix and a right mastering mix based on the intermediate mix. For example, in a possible implementation, a second set of instances of the second network may be configured to take the intermediate mix as input and generate a left mastering mix and a right mastering mix based thereon. As an example, the left mastering mix and the right mastering mix may be generated by summing (e.g., averaging) all left channel signals processed by the second set of second networks and all right channel signals processed by the second set of second networks, respectively. The system may further include a pair of (weight-sharing) instances of the second network, which may be configured to generate output mixes based on the left mastering mix and the right mastering mix.
[0025] In some examples, the first network may be configured to further generate (eg, master control) parameters for the pair of instances of the second network.
[0026] In some examples, the second network may be trained by using any suitable means. Training may refer to determining parameters for a deep learning model (e.g., a neural network) used to implement the system. Further, training may refer to iterative training.
[0027] In a possible implementation, training the second network may include receiving at least one second training set as input. In particular, the second training set may include a plurality of audio signals, and for each audio signal, at least one transformation parameter for signal processing of the audio signal and a respective predetermined processed audio signal. The predetermined processed audio signal may be provided (e.g., processed) by, for example, an audio engineer or by any other suitable means, such that the predetermined processed audio signal may represent an appropriate and acceptable processed audio signal that may serve as a suitable basis (target) for training. The training may further include inputting the second training set to the second network and iteratively training the second network to predict each processed audio signal based on the audio signal and the transformation parameter. In particular, the training may be based on at least one second loss function indicating a difference between the predetermined processed audio signal and each of its predictions.
[0028] In some examples, the parameters generated by the first network may be human- and / or machine-interpretable parameters. Human-interpretable parameters may generally mean that the parameters can be interpreted (understood) by a human, e.g., an audio engineer, so that the audio engineer can (directly) use or apply those parameters for further audio signal processing or analysis if deemed necessary. Similarly, machine-interpretable parameters may generally mean that the parameters can be interpreted by a machine, e.g., a computer or a program stored therein, so that the parameters can (directly) be used by a program (e.g., a mixing console) for further audio signal processing or analysis if deemed necessary.
[0029] In some examples, the parameters generated by the first network may include control parameters and / or panning parameters, as described above. Of course, the parameters generated by the first network may include any other suitable parameters (e.g., parameters that exist in a real-world mixing console) according to various implementations.
[0030] In some examples, the first and / or second network may include at least one neural network, particularly a neural network that may include at least one of a linear layer, a multilayer perceptron (MLP), etc., as will be appreciated by those skilled in the art.
[0031] In some examples, the neural network may be a convolutional neural network (CNN), such as a temporal convolutional network (TCN), or a Wave-U-Net, or a recurrent neural network, or may include an attention layer or a transformer. Of course, any other suitable neural network may be applied, as recognized by those skilled in the art.
[0032] According to another aspect of the present disclosure, a deep learning-based system for automatic multi-track mixing based on multiple input audio tracks is provided. The system may include a transformation network. In some cases, the transformation network may correspond to the second network described above. In particular, the transformation network may be configured to apply signal processing and at least one mixing gain to the input audio tracks to generate an output mix of the audio tracks based on one or more parameters. The parameters may be generated by other network elements (e.g., a controller network or the first network described above).
[0033] In some examples, the parameters may be human-interpretable parameters. For example, the parameters may be interpretable (or usable) by an audio engineer, so that the audio engineer can (directly) use or apply those parameters for further audio signal processing or analysis if deemed necessary. In some cases, human (or machine) interpretable parameters may simply refer to those found on a conventional / real-world mixing console.
[0034] In some examples, a system may have multiple instances of a first network in a weight-sharing configuration and / or multiple instances of a second network in a weight-sharing configuration. As discussed above, a weight-sharing configuration may generally mean that all of the first and / or second networks may be configured (applied) with the same weights (e.g., weight vector).
[0035] According to another aspect of the present disclosure, a method of operating a deep learning-based system for automatic multi-track mixing based on multiple input audio tracks is provided. The system may include one or more instances of a first deep learning-based network and one or more instances of a second deep learning-based network. The method may include generating, by the first network, parameters used for the automatic multi-track mixing based on the input audio tracks. Additionally, the method may further include applying, by the second network, signal processing and at least one mixing gain to the input audio tracks based on the parameters to generate an output mix of the audio tracks.
[0036] According to another aspect of the present disclosure, a method for training a deep learning-based system for automatic multi-track mixing based on multiple input audio tracks is provided. Training may refer to determining parameters for a deep learning model (e.g., a neural network) used to implement the system. Furthermore, training may refer to iterative training. The system may include one or more instances of a first deep learning-based network and one or more instances of a second deep learning-based network. In particular, the method may include a (first) training phase for training the second network, where the (first) training phase for training the second network may include obtaining as input at least one first training set, the first training set including a plurality of audio signals, each of the audio signals including at least one transformation parameter for signal processing of the audio signal, and each predetermined processed audio signal. The (first) training phase for training the second network may further include inputting the first training set to the second network and iteratively training the second network to predict each processed audio signal based on the audio signals and the transformation parameters in the first training set. In particular, the training of the second network may be based on at least one first loss function indicative of the difference between a given processed audio signal and its respective predictions.
[0037] In some examples, the method may further include a (second) training phase for training the first network, which may include receiving at least one second training set as input, the second training set including a plurality of subsets of audio tracks and, for each subset, a predetermined mix of each of the audio tracks in the subset. The (second) training phase for training the first network may further include inputting the second training set to the first network and iteratively training the first network to predict the mix of each of the subset's audio tracks in the second training set. In particular, the training of the first network may be based on at least one second loss function indicative of the difference between the predetermined mix of the audio tracks and their respective predictions.
[0038] In some examples, the (second) training phase of training the first network may begin after the (first) training phase of training the second network has finished, i.e., the training of the first network may be performed using a pre-trained second network.
[0039] According to a further aspect of the present disclosure, there is provided a computer program, which may include instructions that, when executed by a processor, cause the processor to perform all of the steps of the exemplary methods described throughout the present disclosure.
[0040] According to a further aspect, there is provided a computer-readable storage medium, which may store the computer program described above.
[0041] According to yet another aspect, an apparatus is provided that includes a processor and a memory coupled to the processor, the processor may be adapted to cause the apparatus to perform all steps of the exemplary methods described throughout this disclosure.
[0042] It will be understood that system features and method steps may be interchanged in many ways. In particular, details of the disclosed methods can be implemented by a corresponding system, and vice versa, as will be understood by those skilled in the art. Furthermore, any statements made above relating to a method will be understood to apply equally to a corresponding system, and vice versa.
[0043] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings. [Brief explanation of the drawings]
[0044] [Figure 1] FIG. 1 is a schematic block diagram of a system for multi-track mixing according to an embodiment of the present disclosure. [Figure 2] FIG. 1 is a schematic block diagram of a system for multi-track mixing according to an embodiment of the present disclosure. [Figure 3] FIG. 1 is a schematic block diagram of a controller network configuration according to an embodiment of the present disclosure. [Figure 4A] FIG. 1 is a schematic block diagram of a training model for a transformation network according to an embodiment of the present disclosure. [Figure 4B] FIG. 10 is a schematic block diagram of a training model for a transformation network according to another embodiment of the present disclosure. [Figure 5A] FIG. 1 is a schematic block diagram of a system for multi-track mixing according to an embodiment of the present disclosure. [Figure 5B] FIG. 1 is a schematic block diagram of a system for multi-track mixing according to an embodiment of the present disclosure. [Figure 5C] FIG. 1 is a schematic block diagram of a system for multi-track mixing according to an embodiment of the present disclosure. [Figure 5D] FIG. 1 is a schematic block diagram of a system for multi-track mixing according to an embodiment of the present disclosure. [Figure 6] FIG. 1 is a schematic block diagram of a system for multi-track mixing according to another embodiment of the present disclosure. [Figure 7] FIG. 2 is a schematic diagram of a block diagram of a routing component according to an embodiment of the present disclosure. [Figure 8A] FIG. 1 is a schematic diagram illustrating a block diagram of a neural network according to an embodiment of the present disclosure. [Figure 8B] FIG. 1 is a schematic diagram illustrating a block diagram of a neural network according to an embodiment of the present disclosure. [Figure 8C] FIG. 1 is a schematic diagram illustrating a block diagram of a neural network according to an embodiment of the present disclosure. [Figure 9] 1 is a flowchart illustrating an example method of operation of a deep learning-based system for automatic multi-track mixing based on multiple input audio tracks, according to an embodiment of the present disclosure. [Figure 10A] 1 is a flowchart illustrating an example method for training a deep learning-based system for automatic multi-track mixing, according to an embodiment of the present disclosure. [Figure 10B] 1 is a flowchart illustrating an example method for training a deep learning-based system for automatic multi-track mixing, according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0045] The drawings and the following description relate to preferred embodiments by way of example only. It should be noted that, from the following discussion, alternative embodiments of the structures and methods disclosed herein will readily become apparent as viable alternatives that may be implemented without departing from the principles of what is claimed.
[0046] Reference will now be made in detail to several embodiments that are illustrated in the accompanying drawings. It should be noted that, where practical, similar or identical reference numbers may be used in the figures and may indicate similar or identical functionality. The figures depict embodiments of the disclosed system (or method) for illustrative purposes only. Those skilled in the art will readily recognize that alternative embodiments of the structures and methods described herein may be implemented without departing from the principles set forth herein.
[0047] As mentioned above, the success of deep learning in related audio tasks appears to be driving interest in applying these models for automated multi-track mixing systems. Due to the lack of parametric mixing console data (i.e., the set of settings used by audio engineers) and the inability to propagate gradients through the mixing console in the training process, end-to-end models operating directly at the waveform level appear to offer the most feasible option. In this disclosure, we aim to learn the mixing practices of audio engineers by directly observing the audio transformation from original instrument recordings to the final stereo mix.
[0048] Unfortunately, the lack of multi-track data and the large amount of data typically required to train deep learning models makes this approach unlikely to be feasible. To address this, this disclosure presents a hybrid approach in which, generally, a model that learns the general signal processing algorithms of a mixing console is constructed first, followed by a second (smaller) model that learns to control these channel-like sub-modules to produce a mix.
[0049] Broadly speaking, a first model (e.g., the claimed second network) may learn to operate directly on audio signals and emulate signal processing algorithms present in conventional mixing consoles. Because algorithms from conventional mixing consoles (e.g., equalizers, compressors, reverberation) may be accessible, this model may be trained with a virtually unlimited supply of generated examples. This trained model may then be used to construct a complete mixing console by configuring multiple instances. A second, smaller model (sometimes called a controller; e.g., the claimed first network) may be trained to generate a set of control signals (or some other suitable signals / parameters) for these instances to generate a quality mix of the inputs. This type of learning was not possible because not all elements in conventional mixing consoles are discriminable. The currently described formulation allows for direct learning of control signals to generate a mix in the waveform domain.
[0050] Referring to FIG. 1, a (simplified) block diagram of a system 1000 for performing multi-track mixing according to an embodiment of the present disclosure is shown. In particular, the system 1000 is intended to take as input N audio tracks (signals) 1100-1, 1100-2, ..., 1100-N and provide a (meaningful) audio mix 1400 thereof. The input audio tracks 1100-1 through 1100-N are simultaneously provided to a controller network 1200 (or sometimes simply referred to as a controller) for analysis (e.g., pre-processing). Generally speaking, in a possible implementation, the controller network 1200 may take these input audio tracks 1100-1 through 1100-N, extract (relevant) information from these input audio tracks 1100-1 through 1100-N, and generate as output a set of parameters 1250 for each audio track (channel). The parameters 1250 may include control parameters, panning parameters, or any other type of parameters suitable for audio mixing. The mixing process may then be performed by a so-called mixing console 1300. The mixing console 1300 may be operated, for example, by an audio engineer, based on the input audio tracks 1100-1 to 1100-N and the parameters 1250 to generate a final output audio mix 1400.
[0051] Generally speaking, the task of transforming a collection of audio signals into a cohesive mixture requires a deep understanding of disparate technical and creative processes. To perform this task effectively, an audio engineer's professional training includes developing the ability to recognize how to utilize a set of signal processing tools to achieve a set of desired technical and creative goals. Thus, there is interest in systems that can perform this process automatically, in order to provide tools to less-skilled users and reduce the time required for experienced engineers, as well as audio engineers.
[0052] In recent years, deep learning has shown impressive results in many audio tasks previously considered extremely difficult (e.g., speech synthesis, instrument synthesis, and audio source separation). For this reason, there is a general interest in applying these models within the context of methods for automatic multi-track mixing. This system may take a large number of input audio recordings from various sources, process them individually, and then combine them to produce a final mix, just as an audio engineer would.
[0053] Unfortunately, there appear to be many challenges to applying deep learning approaches to this task, and so far, these challenges have prevented such applications from being fully implemented. One of the most significant challenges is the limited size of available training data. This reality also makes it unlikely that standard deep neural networks, whose input is a collection of multi-track recordings and whose output is a mix of those tracks, can be trained in an end-to-end manner. Training an end-to-end model operating in the waveform domain can require more than a million samples to perform effectively on classification tasks, and while spectrogram-based approaches have been shown to perform more competitively with less data, such approaches are even more problematic for tasks that use audio as their output due to their challenges in modeling phase.
[0054] In consideration of some or all of the above challenges, it has been generally proposed to replace the mixing console 1300 itself with a set of neural networks, allowing for the construction of a fully discriminative mixing console equipped with deep learning building blocks. To do so, it should be noted that the mixing console 1300 itself consists of a series of repeated channels, with all channels enabling the same set of transformations, e.g., using processor configurations. Thus, to emulate the mixing console 1300, generally speaking, one need only emulate a single channel in the console and then apply weight sharing among all channels. Roughly speaking, this can be achieved by designing an appropriate network, training it to emulate the signal processing chain of a single channel in a mixing console, and extending this network across multiple input recordings. Ideally, when this network is given an audio signal and parameters for the processors in the channel, it will produce an audio signal that is indistinguishable from the output of a true mixing console channel (e.g., as operated by an audio engineer).
[0055] The above concept is illustrated diagrammatically in FIG. 2. In particular, the same or similar reference numerals in FIG. 2 may refer to the same or similar elements of the system 1000 shown in FIG. 1, so that repeated descriptions thereof may be omitted for the sake of brevity. As shown in FIG. 2, the mixing console 1300 of FIG. 1 is replaced by multiple (neural) networks 2310-1, 2310-2, ..., 2310-N, each of which may be trained to emulate the signal processing chain of a single channel within the mixing console. In this sense, the neural networks may sometimes be referred to as transformation networks. In the example of FIG. 2, the number of neural networks 2310-1 to 2310-N is equal to the number of tracks (channels) of the inputs 2100-1 to 2100-N, although it should be understood that these numbers do not necessarily always have to be the same. 2, each network may be applied across each input channel, with controller network 2200 generating a set of parameters for each input channel to generate a mix that is the sum of the outputs of each network (e.g., using summation component 2350). Notably, this approach (i.e., all N instances of networks 2310-1 through 2310-N are applied the same set of parameters) is sometimes referred to as weight sharing.
[0056] Such a configured proposed design may enable or facilitate training of the controller network, such that the complete system is sufficiently discriminative. Furthermore, the present disclosure also provides the advantage of easily scaling to many different effects (or signal processing) to be applied in a mixing scenario, potentially expanding the size of the signal chain without having to create a unique discriminative implementation of each new digital audio effect. This can significantly reduce the complexity of the design and / or training, and significantly improve the efficiency of the overall system.
[0057] In summary, transformation networks 2310-1 through 2310-N can replace traditional channels in a typical mixing console and attempt to operate in the same (shared weights) manner when fed a set of input signals and parameters. By configuring multiple instances of the same (pre-trained) transformation network, a complete, discriminative "mixing console" 2300 can be constructed, ultimately enabling training of the controller network while facilitating learning from limited data. Notably, in some cases, controller network 2200 of system 2000 may be referred to simply as the "first network," while the transformation network may be referred to simply as the "second network."
[0058] As described above, the training of the controller network (first network) may be performed separately from the training of the transformation network (second network). In some possible implementations, the transformation network may be trained prior to the training of the controller network, or in other words, the training of the controller network may depend on a pre-trained transformation network.
[0059] Figure 3 is a schematic block diagram of a configuration 3000 of a controller network 3200 according to an embodiment of the present disclosure. The controller network 3200 may be similar to that shown in Figure 1 or 2. Thus, broadly speaking, the controller network 3200 may be viewed as having at least one (deep) neural network trained to generate a set of parameters 3250 (e.g., for a mixing console or one or more transformation networks) to produce a desired mix given an input 3100 audio waveform.
[0060] 3, the controller network 3200 may have two stages (or two sub-networks): a first stage 3210 and a second stage 3220. In some cases, the first stage 3210 may be referred to as an encoding stage (or simply an encoder), while the second stage 3220 may be referred to as a post-processing stage (or simply a post-processor).
[0061] Generally speaking, the encoding stage 3210 may be responsible for extracting relevant information from the input channels. For example, such extraction may involve converting (or mapping) the input audio waveform to a feature space representation (e.g., a latent space representation). The type of information that may be relevant to the mixing task may be characteristics such as the source of the input (e.g., guitar, drums, voice, etc.), as well as detailed information such as the energy envelope over time and the allocation of energy across the frequency spectrum, which may be needed to understand masking interactions between sources. Typically, these are the same or similar types of considerations that an audio engineer would make when trying to (manually) create an audio mix. The encoding stage 3210 may then generate a representation of each input signal (channel). For this purpose, multiple encoding stages 3210 may comprise (sub)encoders 3211-1, 3211-2, ..., 3211-N (which may, for example, be equal to the number of channels / tracks in the input 3100). The output of the encoding stage 3210 is then passed to the post-processing stage 3220.
[0062] Broadly speaking, the role of the post-processing stage 3220 may be to aggregate information from the encoders in order to make decisions regarding how to parameterize the transformation network that operates on the associated input recording. In particular, in some instances, such decisions may not be made in isolation from the other input channels, in that each mixing decision may generally depend heavily on some or all of the other inputs. With this in mind, the post-processing stage 3220 may be provided with not only training representations for each input audio track, but also a combined (or concatenated) representation that somehow represents or summarizes some or all of the inputs. In a possible implementation, this may be achieved by a simple average (as exemplified by 3215) over all input representations output by the encoders 3211-1 to 3211-N. Of course, any other suitable means may be adapted to generate an appropriate combined representation depending on various implementations and / or requirements. In some cases, such a combined representation may also be referred to as a context representation. Based on the combined representation, the post-processing stage 3220 may then be configured (trained) to output a set of parameters that can be used for audio mixing (e.g., by a transformation network, or any other suitable network component). Similar to the encoding stage 3210, the post-processing stage 3220 may itself have any number of (sub) post-processors 3221-1, 3221-2, ..., 3221-K, depending on various implementations. In particular, in some possible implementations, the weight sharing concept described above may also be applied (extended) to the controller network 3200. As an example, one instance (pair) of a sub-encoder (e.g., 3211-1) and a post-processor (e.g., 3221-1) may be applied individually to each of the input channels, and a set of parameters for each channel is generated and passed to the transformation network (which, as described above, also shares weights). As such, weight sharing may be considered to be applied at the (full) system level. An example of such system-level weight sharing is also shown in FIG. 5 and will be discussed in more detail below.
[0063] 4A is a schematic diagram of a block diagram of a training model 4100 for a transformation network 4311 according to an embodiment of the present disclosure. The transformation network 4311 can be the same as or similar to the transformation network 2310-1 shown in FIG.
[0064] Generally speaking, the goal of a transformation network is to build a model that can implement common signal processing tools (e.g., equalizers, compressors, reverberation) utilized by audio engineers in a discriminative manner. This network is necessary because traditional signal processing algorithms implement functions with gradients that can be ill-behaved or unwieldy, making them difficult to use in the process of training a model to control them to control a mix. In broad terms, a transformation network can take as input an audio signal and a set of parameters defining the control of all processors in the signal chain, possibly along with their respective orderings. The transformation network can then be trained to produce the same output as the true set of signal processors. During training, input / output pairs can be generated, for example, by randomizing the states of processor parameters in the true signal chain and then passing various signals through this signal chain to generate target waveforms. These pairs are then used to train the transformation network. In this sense, such a training process can be considered to be completed in a self-supervised manner, using a nearly infinite set of training data. The training data may then be collectively compiled into at least one training set that is used to (iteratively) train the transformation network, as will be appreciated by those skilled in the art.
[0065] 4A, a transformation network 4311 takes as input two sets of values: the input audio waveform 4101 itself and corresponding parameters 4251 for controlling signal processing of the audio waveform 4101. In some cases, the signal processing may relate to audio effects such as (but not limited to) gain, panning, equalization, dynamic range compression, reverberation, etc., as will be understood by those skilled in the art. The input audio 4101 and parameters 4251 are input to both the network to be trained 4311 and a conventional mixing console 4301 so that the outputs of each, i.e., the "true" processed audio signal 4421 and the predicted processed audio signal 4411, can be compared to facilitate further training processes.
[0066] 4B schematically depicts a block diagram of a (slightly more detailed) training model 4200 of a transformation network 4312 according to another embodiment of the present disclosure. In particular, the same or similar reference numerals of the model 4200 of FIG. 4B indicate the same or similar elements of the model 4100 shown in FIG. 4A, whereby repeated description thereof may be omitted for brevity. As above, broadly speaking, the transformation network 4312 may be generally considered to comprise at least one deep neural network trained to mimic the processing of conventional channels in a mixing console, given a set of parameters defining the configuration of the channels.
[0067] 4B , the transformation network 4312 is shown to have two sub-networks 4313 and 4314, both of which may be neural networks themselves. In possible implementations, the first sub-network 4313 may be a parametric deep neural network (or P-DNN for short) that takes as input the input audio waveform 4102 and parameters 4252, while the second sub-network 4314 may be a transformation deep neural network (or T-DNN for short) that takes as input the input audio 4102 and the results of the P-DNN 4313, and may output a prediction 4412 of the processed audio signal. In some possible implementations, training may be performed using a loss function 4502 that may indicate the difference between a given (“true”) processed audio signal 4422 (as processed by the mixing console channel 4302 based on the audio input 4102 and the corresponding parameters 4252) and their respective predictions 4412.
[0068] Using this trained transformation network according to either Figure 4A or Figure 4B, a differential mixing console proxy can then be constructed, which can utilize multiple instances of these pre-trained transformation networks. As described above, weights can be shared between all instances, so that processing on each channel of the console functions identically. A controller network can then be introduced, whose purpose is generally to generate parameter adjustments for each individual transformation network given information about the input audio channels.
[0069] Broadly speaking, to train a controller network (e.g., controller network 3200 shown in FIG. 3), paired multi-track stems may be fed as input to a transformation network instance, and a mix may be created by the controller network. The mix may then be compared to a ground truth mix to train the controller network. This is made possible because all elements are fully discriminable. In some possible implementations, training may be performed based on at least one loss function that indicates the difference between a given mix of audio tracks (the "ground truth mix") and their respective predictions. Optionally, during the training process, the weights of the controller network may be fine-tuned. Finally, the proposed system may generally enable the ability to learn directly at the waveform level from a limited set of multi-track stems and their respective mixes.
[0070] A possible implementation of a complete system 5000, i.e., both the controller network and the transformation network, is shown schematically in Figure 5D. In particular, the system 5000 may have a controller network (also referred to as a first network) 5200 that takes as input a number of audio signals 5100-1, 5100-2, ..., 5100-N and generates a set of parameters (e.g., control parameters, panning parameters, etc.). The input audio tracks 5100-1 to 5100-N, along with the generated parameters, are then fed to one or more instances of a transformation network (also referred to as a second network) 5300 to transform the final output mix 5400. More specifically, the role of the controller network 5200 is to process the input channels 5100-1 to 5100-N, extract useful information about those inputs (e.g., by an encoding stage 5210), process the extracted information (possibly based on a combined representation illustrated as 5215) (e.g., by a post-processing stage 5220), and finally generate a set of parameters that are sent to the associated transformation network 5300 to process the input channel in question.
[0071] As mentioned above, in the example of Figure 5D, weight sharing is applied at the system level. That is, a pair of encoders and post-processors is applied individually to each of the input channels to generate a set of parameters for each channel, which are then passed to a set of transformation networks that also share weights. As such, overall system efficiency can be further improved. Nevertheless, it should be understood that any other suitable configuration (of the controller network 5200 and / or transformation network 5300) may be used, as will be understood by those skilled in the art.
[0072] 5D, output mix 5400 is generated as the sum (illustrated as 5350) of the outputs of transformation networks 5300. However, it should be understood that any other suitable operations / processes may be applied to generate the final mix. As a general example (but not by way of limitation), at least one mixing gain (e.g., depending on whether a stereo output is desired) may simply be applied to generate the audio mix. In some other possible examples, one mixing gain and one panning parameter may be employed, which may simply imply that the parameters of the other channels correspond to one minus such panning parameter.
[0073] Additionally, for the sake of completeness, Figures 5A-C respectively schematically represent possible implementations of an instance of the encoding stage 5210, an instance of the post-processing stage 5220, and an instance of the transformation network 5300. Specifically, in the example shown in Figure 5A, the encoder may have a series of convolutional blocks and fully connected (FC) blocks (which themselves may have, for example, a series of linear layers and a rectified linear unit (ReLU) activation function). In the example of Figure 5B, the post-processor may have a series of fully connected subnetworks, each of which may have a number of hidden linear layers, a parametric ReLU (PReLU) activation, and a dropout (e.g., p = 0.1), followed by an output layer. Notably, as described above, in addition to each latent space representation generated by the encoding stage, the post-processor (or post-processing stage in general) may optionally further take as input a joint (context) representation so that information embedded in (some or all) of the input channels can be captured. 5C, the transformation network may have a series of fully connected sub-networks, each of which may have an FC block, a series of temporal convolutional networks (TCNs, discussed in more detail below), and an output layer. Of course, as will be understood by those skilled in the art, these implementations are merely exemplary, and any other suitable implementations may be employed.
[0074] 6 is a schematic block diagram of a system 6000 for multi-track mixing according to another embodiment of the present disclosure. In particular, system 6000 may be considered an enhanced version of system 5000, i.e., with several possible enhancements that are likely to be present in a (more complete) process of audio mixing. Thus, the same or similar reference numerals of system 6000 in FIG. 6 may still refer to the same or similar elements of system 5000 shown in FIG. 5D, whereby repeated descriptions thereof may be omitted for the sake of brevity.
[0075] Broadly speaking, complete system diagram 6000 shows the configuration of multiple elements that make up a complete, discriminative mixing console and controller network 6200. Inputs 6100 to system 6000 may be a collection of (e.g., N) audio signals that may correspond, for example, to recordings of different instruments. In some possible implementations, each channel may take a single mono input signal as input for processing. Initially, these inputs 6100 are passed to controller network 6200, which may perform analysis and then generate (e.g., control) parameters 6250 for some or all of the processing elements in system 6000. Those inputs 6100, along with corresponding generated (e.g., channel control) parameters 6251, are then passed to (a first set of) a series of transformation networks 6310-1, 6310-2, ..., 6310-N. The result of this processing (the first phase / stage of mixing) is an output that can be a stereo mix of the tracks (illustrated by the two outputs of each transformation network). As mentioned above, multiple signal processing elements can be included in the system's signal path to simulate a traditional mixing console used by an audio engineer. Referring to the example of FIG. 6, system 6000 can consist of N input channels, M stereo buses, and a master bus, which can generally follow the possible structure of a traditional mixing console. Compared to a traditional mixing console, the main difference is that these network elements can be implemented by neural networks in this case and are therefore discriminative. This makes it possible to train the system to mimic the actions of an audio engineer in an efficient and flexible manner by training the system on a series of mixes created by the audio engineer.
[0076] More specifically, first, each of the N inputs 6100 may be passed to a controller network 6200. This controller network 6200 performs an analysis of the input signals 6100 in an attempt to understand how the inputs should be processed to produce a desired mix. To accomplish this, as described above, a trained encoder may be used to generate a condensed representation of the inputs, extracting the most important information. The extracted information is then used by a neural network (e.g., a post-processor with linear layers and non-linear activations) to make a prediction of how parameters 6250 for the entire mixing console should be set to achieve the desired mix of inputs. As described above, the controller network may further include combining (or, in some cases, concatenating) these feature space (condensed) representations to generate a complete (combined) representation of the inputs, and parameter prediction may then be further performed based on such a combined representation.
[0077] Those parameters 6250 are then passed to the transformation network, which can then begin multi-stage processing of the input signal.
[0078] According to their predicted parameters 6251, each of the N inputs 6100 is passed through several instances (which may be referred to as a first set) of transformation networks 6310-1 to 6310-N. In the example of FIG. 6, each of the transformation networks is shown having two subelements, for example, a T-DNN and a P-DNN, as illustrated in FIG. 4B. The transformation networks may be pre-trained on another task before being used in the complete system 6000, as described in detail above with reference to FIGS. 4A and 4B. In particular, the training process is generally a self-supervised training task, where the input may be a mono waveform and a set of parameters that fully define the configuration of an actual mixing console channel. The output may be a stereo waveform intended to predict the output produced by the actual mixing console channel. Because the P-DNN and T-DNN are each neural networks that can be composed of various architectures, their configurations are not limited in this system design. After this pre-training process, the transformation network can closely mimic a real mixing console channel, with the advantage that all of its elements are discriminable, allowing for further training.
[0079] Because the transformation networks 6310-1 to 6310-N are trained, when passed a set of parameters 6251 and an output 6100, the transformation networks 6310-1 to 6310-N (ideally) perform processing similar to that performed by actual channels in a mixing console, given the same inputs. In the example of FIG. 6, each of the N input channels 6100 may generate a stereo output. These outputs may then be sent to a routing component or subsystem (sometimes simply referred to as a router) 6500. The router 6500 may exist for the purpose of enabling the generation of bus-level mixes for a second (sub)stage of processing. In other words, the router 6500 may be considered a simple mixer that generates submixes and sends them (i.e., the submixes) for further processing, for example, in the console or by a transformation network. To operate such a routing component 6500, the controller network 6200 may be configured to further generate corresponding (e.g., router control) parameters 6252.
[0080] A possible (slightly more detailed) implementation of the routing component 6500 is shown schematically in FIG. 7. Specifically, within the router 7500, parameters 7252 (e.g., provided by the controller 6200 as shown in FIG. 6) are used to generate multiple (e.g., M) unique bus level mixes 7520-1 through 7520-M from N input channel outputs 7510-1, 7510-2, ..., 7510-N. Generally speaking, each input channel can be considered a channel that takes a single mono input signal as input for processing, while a bus may in some cases correspond to a channel that takes the sum of multiple stereo input signals as input for further processing. In some possible implementations, the routing component 7500 can also route channel outputs directly to master buses (left and right), as illustrated in FIG. 7 as 7530-1 and 7530-2. Specifically, a master bus can generally refer to the bus that collects all signals in the system and generates the final stereo output mix.
[0081] In summary, the router 7500 may generally be considered a component (subsystem) that processes the signal flow for the second round of processing. The second round of processing may include generating bus-level mixes for the M buses using predicted parameters 7252 (e.g., from the controller network) and transmitting those mixes. Optionally, the router may also transmit copies of each original input to the left and right master buses.
[0082] Referring again to the example of FIG. 6, the bus level mix output by router 6500 may then be sent to multiple (e.g., M) stereo-linked transformation network pairs 6320-1 (consisting of transformation networks 6320-1A and 6320-B), along with additional (e.g., bus control) parameters 6253 also generated by controller network 6200. Stereo-linked transformation network pair 6320-1 forms M buses in system 600 as shown in FIG. 6. In some cases, these stereo-linked transformation network pairs 6320-1 through 6320-M may simply be referred to as a second set of transformation network instances. These buses may then each generate a dual stereo output, which are summed (e.g., by 6350-1 and 6350-2, respectively) to form a single stereo output. The single stereo output is then sent to a master bus. Finally, the master bus may form final (e.g., stereo-linked) transformation networks 6330-1 and 6330-2, which generate the final output mix 6400 based on further (e.g., master control) parameters 6254 generated by the controller network 6200. In particular, a stereo-linked connection may generally refer to the configuration of a pair (two) transformation networks, each with its own input, where both transformation networks use the same parameters to apply the same signal processing to each input and each generate a stereo output. Furthermore, dual stereo may generally refer to the configuration of a system where the input is a stereo signal and the output provides separate signals for left and right outputs (channels) that are later summed to generate an overall left and right stereo signal. Such configurations / connections (i.e., stereo-linked and dual stereo) are also clearly shown in the example of FIG. 6. In particular, when a stereo output mix is generated, the controller network may, in some implementations, be trained to be invariant under left and right channel reassignments.In some possible implementations, such invariance between stereo channels may be achieved by considering the sum of the audio signals corresponding to the left and right channels, rather than considering them separately.
[0083] In summary, to address at least some or all of the above-mentioned problems, the present disclosure seeks to design a system having two networks: a first network (a transformation network) is pre-trained in a self-supervised manner to instill domain knowledge, and then a second (smaller) network (a controller network) is trained using limited multi-track data to most effectively control the operation of a set of instances of the first network to produce a high-quality mix. Multiple instances of this first network (the transformation network) may be used to construct a system that mimics the design of a traditional mixing console, in that it has multiple channels, routable buses, and a single additive master through which all signals are routed. In general, the goal, at least in part, is to design a second network (a controller network) that can learn to control the transformation networks by learning from a limited data set of paired multi-track stems and mixes. In particular, similar to the (channel control) parameters 6251 above, the (router control) parameters 6252, (bus control) parameters 6253, and / or (master control) parameters 6254 may be generated (e.g., predicted) based on the input audio signals, optionally even based on a combined (or concatenated) representation of all input audio signals. Furthermore, it should be noted that since a general goal of training a transformation network is to mimic a mixing console, the parameters 6250 generated by the controller network can be human- and / or machine-interpretable parameters. Human-interpretable parameters may generally mean that the parameters can be interpreted (understood) by a human, for example, an audio engineer, so that the audio engineer can (directly) use or apply those parameters for further audio signal processing or analysis if deemed necessary.Similarly, machine-interpretable parameters may generally mean that the parameters can be interpreted by a machine, e.g., a computer or a program stored therein, so that they can be used (directly) by a program (e.g., a mixing console) for further audio signal processing or analysis, if deemed necessary. Providing interpretable parameters in this manner enables user interaction, such as adjusting the output mix as needed. Furthermore, such interpretability also allows user interaction to easily fine-tune the training model according to their goals, and the predictions accordingly, thereby further improving the overall system performance while maintaining sufficient flexibility (e.g., in terms of further adjustments or fine-tuning, if needed).
[0084] Furthermore, although only one controller network may be present in the examples shown in Figures 1, 2, 5, and 6, it should nevertheless be understood that in some other cases, more than one instance of a controller network may be provided in the system as well. For example, in some possible implementations, different instances of a controller network may be provided for each of the conversion networks, routing components, or some of the master buses. That is, different controller networks may be provided to generate parameters that can each be used by different sub-networks.
[0085] 8A-8C schematically illustrate some possible implementations of neural networks according to some embodiments of the present disclosure. However, it is worth noting that, as will be understood by those skilled in the art, these possible implementations should be understood as merely examples, and any other suitable form of neural network may be implemented.
[0086] In particular, FIG. 8A may represent a schematic diagram of a high-level view of a TCN architecture 8100 (e.g., suitable for implementing the transformation network shown in FIG. 5C). Generally speaking, a TCN can be viewed as formulating the application of a convolutional neural network (CNN) to time-series data and may include many components, such as one-dimensional kernels, convolutions with exponentially increasing expansion factors, and residual connections. In the exemplary architecture shown in FIG. 8A, a stack (e.g., 10) of convolution blocks (denoted as TCNn) with possibly exponentially increasing expansion factors is provided. Each block in the stack may have a residual connection and an additional skip connection to the output. Feature-wise linear modulation (or FiLM for short) generally refers to a conditioning method that formulates the technique as a learned affine transformation performed on the intermediate features of a CNN.
[0087] An example of a possible implementation of a single convolution block 8200 (denoted TCNn in FIG. 8A) is shown schematically in FIG. 8. Specifically, the convolution block may consist of a standard formulation featuring one-dimensional convolution, batch normalization, an affine transformation to insert conditioning via the FiLM mechanism, and finally PReLU. To calculate the final output of the block, a residual connection with a learned scaling factor (denoted gn in FIG. 8B) is included.
[0088] 8C schematically depicts a high-level view 8300 of another possible implementation of a neural network, which in this case may be an architecture based on Wave-U-Net. Generally speaking, Wave-U-NET can be viewed as an adaptation of the traditional U-Net architecture to operate on waveforms, including some additional approaches to incorporate additional input context, possibly using strided transposed convolutions for upsampling.
[0089] Nevertheless, as mentioned above, these possible implementations of neural networks may merely serve for illustrative purposes: any other suitable form, such as a recurrent neural network (RNN), or including an attention layer or a transformer, may likewise be employed.
[0090] 9 is a flowchart illustrating an example method 900 of operation of a deep learning-based system for automatic multi-track mixing based on multiple input audio tracks, according to an embodiment of the present disclosure. The system may be the same as or similar to, for example, system 5000 shown in FIG. 5 or system 6000 shown in FIG. 6. That is, the system may include one or more instances of a suitable controller network (or simply referred to as a first network) and one or more (weight-sharing) instances of a suitable transformation network (or simply referred to as a second network), as shown in either figure. Accordingly, repeated descriptions thereof may be omitted for the sake of brevity.
[0091] In particular, the method 900 may begin at step S910 with generating, by a first network, parameters to be used for automatic multi-track mixing based on the input audio tracks.
[0092] Thereafter, the method 900 may continue with step S920 of applying, by a second network, signal processing and at least one mixing gain to the input audio tracks based on the parameters to generate an output mix of the audio tracks.
[0093] 10 is a flowchart illustrating an example method for training a deep learning-based system for automatic multi-track mixing, according to an embodiment of the present disclosure. The system may be the same as or similar to, for example, system 5000 shown in FIG. 5 or system 6000 shown in FIG. 6. That is, the system may include one or more instances of a suitable controller network (or simply referred to as a first network) and one or more (weight-sharing) instances of a suitable transformation network (or simply referred to as a second network), as shown in either figure. Accordingly, repeated descriptions thereof may be omitted for the sake of brevity.
[0094] As mentioned above, training of the second network (i.e., the transformation network) may be performed before that of the first network (i.e., the controller network). Thus, broadly speaking, training of the entire system can be considered as being divided into two separate training phases. The two training phases are illustrated in Figures 10A and 10B, respectively.
[0095] 10A schematically illustrates an example of a (first) training phase 1010 for training a second network (i.e., a transformation network), starting with step S1011 of obtaining at least one first training set as input. The first training set may include a plurality of audio signals, and for each audio signal, may include at least one transformation parameter for signal processing of that audio signal and each predetermined processed audio signal. The training phase 1010 then continues with step S1012 of inputting the first training set to the second network, followed by step S1013 of iteratively training the second network to predict each processed audio signal based on the audio signals and the transformation parameters in the first training set. More specifically, the training of the second network may be based on at least one first loss function indicative of the difference between the predetermined processed audio signal and each of its predictions.
[0096] FIG. 10B further illustrates a schematic example of a (second) training phase 1020 for training a first network (i.e., a controller network), starting with step S1021 of receiving at least one second training set as input. The second training set may include multiple subsets of audio tracks, and for each subset, may include a predetermined mix of each of the audio tracks in that subset. The training phase 1020 then continues with step S1022 of inputting the second training set to the first network, followed by step S1023 of iteratively training the first network to predict the mix of each of the subset's audio tracks in the second training set. More specifically, the training of the first network may be based on at least one second loss function indicative of the difference between the predetermined mix of audio tracks and their respective predictions. In some possible cases, the predetermined mix of audio tracks may be a stereo mix. Accordingly, the second loss function may be a stereo loss function and may be configured to be invariant under left and right channel reassignments. In some possible implementations, such invariance between stereo channels may be achieved by considering the sum of the audio signals corresponding to the left and right channels, rather than considering them separately.
[0097] In particular, the present disclosure may be utilized in several ways in which automatic mixing may be important. For example (but not by way of limitation), in some use cases, a user may upload separated / instrumental tracks, or they may be obtained from a music source separation algorithm. An automatic mixing process may then take place, which may improve the quality of the user-generated content and become a distinctive feature of the product. Another potential opportunity is when a recording engineer starts from an initial mix provided by the proposed approach, which may further include spatial mixing functions. Of course, any other suitable use case may be utilized, as will be understood and appreciated by those skilled in the art. Furthermore, another possibility may be when a user provides (e.g., uploads) an already mixed audio signal. Then, if there is some source separation algorithm that can decompose the mix into separate tracks, these separate tracks can be auto-mixed (again) by using the approach of the present disclosure. The result will be a different mix based on the mix signal and may also include human intervention to refine the automix result before generating the final mix.
[0098] Thus far, possible training and operation methods for a deep learning-based system for determining an indication of audio quality of an input audio sample, as well as possible implementations of such a system, have been described. Additionally, the present disclosure also relates to apparatus for performing these methods. Examples of such apparatus may include a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), one or more application specific integrated circuits (ASICs), one or more radio frequency integrated circuits (RFICs), or any combination thereof) and a memory coupled to the processor. The processor may be adapted to perform some or all of the steps of the methods described throughout this disclosure.
[0099] An apparatus may be a server computer, a client component, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a smartphone, a web appliance, a network router, a switch or bridge, or any machine capable of executing instructions (sequential or otherwise) that specify operations to be performed by the apparatus. Furthermore, this disclosure relates to any collection of apparatuses that individually or collectively execute instructions to perform any one or more of the methodologies discussed herein.
[0100] The present disclosure further relates to a program (e.g., a computer program) having instructions that, when executed by a processor, cause the processor to perform some or all of the steps of the methods described herein.
[0101] Furthermore, the present disclosure relates to a computer-readable (or machine-readable) storage medium that stores the above-described program, where the term "computer-readable storage medium" includes, but is not limited to, data repositories in the form of solid-state memory, optical media, and magnetic media.
[0102] Unless otherwise stated, and as will be apparent from the discussion that follows, throughout this disclosure, discussions using terms such as "processing," "computing," "calculating," "determining," "analyzing," and the like will be understood to refer to the operations and / or processes of a computer or computing system or similar electronic computing device that manipulates and / or transforms data represented as physical, e.g., electrical, quantities into other data similarly represented as physical quantities.
[0103] Similarly, the term "processor" can refer to any device or part of a device that processes electronic data, e.g., from registers and / or memory, and converts the electronic data into other electronic data, which may be stored, e.g., in registers and / or memory. A "computer" or "computing machine" or "computing platform" may include one or more processors.
[0104] In one exemplary embodiment, the methodologies described herein are executable by one or more processors that accept computer-readable (also called machine-readable) code, which includes a set of instructions that, when executed by the one or more processors, perform at least one of the methods described herein. Any processor capable of executing a set of instructions that specify operations to be performed is included. Thus, one example is a typical processing system including one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system may further include a memory subsystem including main RAM and / or static RAM and / or ROM. A bus subsystem may be included for communication between components. The processing system may be a distributed processing system in which processors are coupled by a network. If the processing system requires a display, such a display may be included, for example, a liquid crystal display (LCD) or a cathode ray tube (CRT) display. If manual data entry is required, the processing system also includes input devices, such as one or more of an alphanumeric input unit such as a keyboard, a pointing control device such as a mouse, etc. The processing system may also include a storage system, such as a disk drive unit. In some configurations, the processing system may also include an audio output device and a network interface device. Thus, the memory subsystem includes a computer-readable carrier medium that carries computer-readable code (e.g., software) that includes a set of instructions that, when executed by one or more processors, cause the execution of one or more of the methods described herein. It should be noted that, when a method includes several elements, e.g., several steps, no order of such elements is implied unless specifically stated. The software may reside on a hard disk, or may reside, completely or at least partially, in RAM and / or within the processor during its execution by the computer system.Thus, the memory and processor also constitute a computer-readable carrier medium carrying computer-readable code. Further, the computer-readable carrier medium may form, or be included in, a computer program product.
[0105] In alternative exemplary embodiments, one or more processors may operate as stand-alone devices or may be connected in a networked arrangement, e.g., networked to other processors, one or more processors may operate as server or user machines in a server-user network environment, or as peer machines in a peer-to-peer or distributed network environment. One or more processors may form a personal computer (PC), tablet PC, personal digital assistant (PDA), cellular telephone, web appliance, network router, switch, or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) resulting in actions being performed by the machine.
[0106] It should be noted that the term "machine" should also be considered to include any collection of machines that individually or collectively execute instructions (set or sets) to perform any one or more of the methodologies discussed herein.
[0107] Thus, one exemplary embodiment of each of the methods described herein takes the form of a computer-readable carrier medium carrying a set of instructions, e.g., a computer program, executed on one or more processors, e.g., one or more processors that are part of a web server arrangement. Thus, as will be appreciated by those skilled in the art, exemplary embodiments of the present disclosure may be embodied as a method, an apparatus such as a dedicated machine, an apparatus such as a data processing system, or a computer-readable carrier medium, e.g., a computer program product. The computer-readable carrier medium carries computer-readable code, which, when executed on one or more processors, causes the one or more processors to perform the method. Thus, aspects of the present disclosure may take the form of a method, an exemplary entirely hardware embodiment, an exemplary entirely software embodiment, or an exemplary embodiment combining software and hardware aspects. Furthermore, the present disclosure may take the form of a carrier medium (e.g., a computer program product on a computer-readable storage medium) carrying computer-readable program code embodied in the medium.
[0108] Software may also be transmitted or received over a network via a network interface device. While the carrier medium is a single medium in the exemplary embodiment, the term "carrier medium" should be understood to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more sets of instructions. The term "carrier medium" should also be understood to include any medium that can store, encode, or carry a set of instructions for execution by one or more processors, causing the one or more processors to perform any one or more of the methodologies disclosed herein. Carrier media can take many forms, including, but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical disks, magnetic disks, and optical-magnetic disks. Volatile media include dynamic memory, such as main memory. Transmission media include coaxial cables, copper wire, and optical fiber, including wiring that comprises a bus subsystem. Transmission media may also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications. For example, the term "carrier medium" should be understood accordingly to include, but not be limited to, solid-state memory, computer products embodied in optical and magnetic media, media having propagated signals detectable by at least one processor or one or more processors and representing a set of instructions that, when executed, perform a method, and transmission media within a network having propagated signals detectable by at least one processor of one or more processors and representing a set of instructions.
[0109] It is understood that the steps of the methods discussed are, in one exemplary embodiment, performed by a suitable processor (or processors) of a processing (e.g., computer) system executing instructions (computer-readable code) stored on a memory device. It is also understood that the present disclosure is not limited to any particular implementation or programming technique, and that the present disclosure may be implemented using any suitable technique that implements the functions described herein. The present disclosure is not limited to any particular programming language or operating system.
[0110] References in this disclosure to "one exemplary embodiment," "some exemplary embodiments," or "exemplary embodiments" mean that a particular feature, structure, or characteristic described in connection with an exemplary embodiment is included in at least one exemplary embodiment of this disclosure. Thus, the appearances of "in one exemplary embodiment," "in some exemplary embodiments," or "in exemplary embodiments" in various places throughout this disclosure are not necessarily all referring to the same exemplary embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner, as would be apparent to one of ordinary skill in the art from this disclosure, in one or more exemplary embodiments.
[0111] As used herein, unless otherwise specified, the use of ordinal adjectives "first," "second," "third," etc. to describe a common object merely means that different instances of the same object are being referred to, and is not intended to imply that the objects so described must be in a given order temporally, spatially, ranked, or in any other way.
[0112] In the following claims and the description herein, any one of the terms "comprising," "comprised of," or "which comprises" is an open-ended term meaning the inclusion of at least the elements / features that follow, but not the exclusion of others. Thus, when used in a claim, the term "comprising" should not be interpreted as being limited to the means, elements, or steps listed thereafter. For example, the scope of the expression "a device having A and B" should not be limited to a device consisting only of elements A and B. As used herein, any one of the terms "including," "which includes," or "that includes" is also an open-ended term meaning the inclusion of at least the elements / features that follow, but not the exclusion of others. Thus, "comprising" is synonymous with "having" and means the same.
[0113] It should be understood that in the foregoing description of exemplary embodiments of the present disclosure, various features of the present disclosure may be grouped together in a single embodiment, figure, or description to simplify the disclosure and aid in understanding one or more of the various inventive aspects. However, this method of disclosure should not be interpreted as reflecting an intention that the claims require more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single, foregoing disclosed exemplary embodiment. Accordingly, the claims following this specification are hereby expressly incorporated into this specification, with each claim standing on its own as a separate exemplary embodiment of the present disclosure.
[0114] Furthermore, while some exemplary embodiments described herein include some features included in other exemplary embodiments but not others, combinations of features of different exemplary embodiments are intended to form different exemplary embodiments within the scope of the disclosure, as will be understood by those skilled in the art. For example, in the following claims, any of the claimed exemplary embodiments can be used in any combination.
[0115] In the description provided herein, numerous specific details are set forth. However, it will be understood that exemplary embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail in order not to obscure an understanding of this description.
[0116] Thus, while what is believed to be the best mode of the disclosure has been described, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the disclosure, and all such modifications and variations are intended to fall within the scope of the disclosure. For example, any formulas described above are merely representative of procedures that may be used. Functions may be added or deleted from block diagrams, and operations may be interchanged between functional blocks. Steps may be added or deleted to methods described within the scope of the disclosure.
[0117] Enumerated exemplary embodiments ("EEE") of the present disclosure have been described above in connection with methods and systems for determining an indication of audio quality of an audio input. Accordingly, embodiments of the present invention may relate to one or more of the following enumerated examples: EEE1. A system for automatic multi-tracking in the waveform domain, comprising: a controller configured to analyze a plurality of input waveforms using a neural network to determine at least one parameter for a plurality of transformation networks and a router; a first transformation network configured to generate a stereo output for each input waveform based on the at least one parameter; a router configured to generate a mix of the stereo output for the plurality of input waveforms and to input the stereo output to a plurality of buses; a second transformation network configured to generate a final stereo output from the outputs of the plurality of buses; A system having:
[0118] EEE2. A deep learning based system for automatic multi-track mixing based on multiple input audio tracks, comprising: one or more instances of a first network based on deep learning; one or more instances of a second network based on deep learning; and the first network is configured to generate parameters used for automatic multi-track mixing based on the input audio tracks; the second network is configured to apply signal processing and at least one mixing gain to the input audio tracks based on the parameters to generate an output mix of the audio tracks. system.
[0119] EEE3. The output mix is a stereo mix. The system described in EEE2.
[0120] EEE4. The first network and the second network are trained separately; the first network is trained based on the pre-trained second network; A system according to EEE2 or 3.
[0121] EEE5. The number of instances of the first network and / or the number of instances of the second network is determined according to the number of input audio tracks. A system according to any one of EEE2 to EEE4.
[0122] EEE6. The first network has a first stage and a second stage; generating the parameters by the first network, said first step comprising mapping each of said input audio tracks to a respective feature space representation; generating parameters for use by the second network based on the feature space representation; having A system according to any one of EEE2 to EEE5.
[0123] EEE7. The second step of generating the parameters to be used by the second network comprises: generating a combined representation based on the feature space representation of the input audio track; generating parameters for use by the second network based on the combined representation; and having The system described in EEE6.
[0124] EEE8. Generating the combined representation includes an averaging process on the feature space representation of the input audio track. The system described in EEE7.
[0125] EEE9. The first network is trained based on at least one loss function indicative of the difference between a given mix of audio tracks and their respective predictions. A system according to any one of EEE2 to EEE8.
[0126] EEE10. The first network obtaining at least one first training set as input, said first training set including a plurality of subsets of audio tracks and, for each subset, a predetermined mix of each of the audio tracks in said subset; inputting the first training set into the first network; iteratively training the first network to predict the mix of each of the subset of audio tracks in the training set; Trained by the training is based on at least one first loss function indicative of the difference between a given mix of the audio tracks and their respective predictions; 10. A system according to any one of EEE2 to EEE9.
[0127] EEE11. The predicted mix of the audio track is a stereo mix; the first loss function is a stereo loss function and is configured to be invariant under left and right channel reassignment; The system described in EEE10.
[0128] EEE12. Training the first network to predict the mix of the audio tracks includes, for each subset of audio tracks: generating, by the first network, a plurality of prediction parameters according to the subset of audio tracks; providing the prediction parameters to the second network; generating, by the second network, a prediction of a mix of the subset of audio tracks based on the prediction parameters and the subset of audio tracks; having 1. A system according to claim 10 or 11.
[0129] EEE13. The number of instances of the second network is equal to the number of input audio tracks; the second network is configured to perform signal processing on each input audio track based at least in part on the parameters to generate a respective processed output; the processed output includes a left channel and a right channel; the output mix is generated based on the processed output. 10. The system according to claim 8, wherein said system is a system according to any one of claims 1 to 8.
[0130] EEE14. The system further comprises a routing component; the routing component is configured to generate a plurality of intermediate stereo mixes based on the processed output; the output mix is generated based on the intermediate stereo mix. The system described in EEE13.
[0131] EEE15. The first network is further configured to generate parameters for the routing component. The system described in EEE14.
[0132] EEE16. The one or more instances of the second network are a first set of one or more instances of the second network; the system further comprising a second set of one or more instances of the second network; the number of instances in the second set of one or more instances of the second network is determined according to the number of intermediate stereo mixes. 1. A system according to claim 14 or 15.
[0133] EEE17. The first network is further configured to generate parameters for the second set of one or more instances of the second network. The system described in EEE16.
[0134] EEE18. The system is configured to further generate a left mastering mix and a right mastering mix based on the intermediate stereo mix; The system further includes a pair of instances of the second network; the pair of instances of the second network are configured to generate the output mix based on the left mastering mix and the right mastering mix. 10. The system according to claim 8, wherein said system is a system according to claim 16 or 17.
[0135] EEE19. The first network is further configured to generate parameters for the pair of instances of the second network. The system described in EEE18.
[0136] EEE20. The second network obtaining at least one second training set as input, said second training set including, for each audio signal, at least one transformation parameter for signal processing of said audio signal and a respective predetermined processed audio signal; inputting the second training set into the second network; iteratively training the second network to predict each processed audio signal based on the audio signal and the transformation parameters; Trained by the training is based on at least one second loss function indicative of a difference between the predetermined processed audio signal and each of its predictions; 10. A system according to any one of EEE2 to EEE119.
[0137] EEE21. The parameters generated by the first network are human and / or machine interpretable parameters. 10. A system according to any one of EEE2 to EEE20.
[0138] EEE22. The parameters generated by the first network include control parameters and / or panning parameters. 10. A system according to any one of EEE2 to EEE21.
[0139] EEE23. The first network and / or the second network comprises at least one neural network; The neural network has linear layers and / or multi-layer perceptrons (MLPs), 10. A system according to any one of EEE2 to EEE22.
[0140] EEE24. The neural network is a convolutional neural network (CNN), such as an inverse convolutional network (ICN) or Wave-U-Net, a recurrent neural network, or includes an attention layer or a transformer. The system described in EEE23.
[0141] EEE25. A deep learning based system for automatic multi-track mixing based on multiple input audio tracks, comprising: A conversion network is provided. the transformation network is configured to apply signal processing and at least one mixing gain to the input audio tracks to generate an output mix of the input audio tracks based on one or more parameters; system.
[0142] EEE26. The parameters are human-interpretable parameters. The system described in EEE25.
[0143] EEE27. The system includes multiple instances of the first network in a weight-sharing configuration and / or multiple instances of the second network in a weight-sharing configuration. 10. A system according to any one of EEE2 to EEE27.
[0144] EEE28. A method of operating a deep learning based system for automatic multi-track mixing based on a plurality of input audio tracks, the system having one or more instances of a first deep learning based network and one or more instances of a second deep learning based network, the method comprising: generating, by the first network, parameters to be used for the automatic multi-track mixing based on the input audio tracks; applying, by said second network, signal processing and at least one mixing gain to said input audio tracks based on said parameters to generate an output mix of said audio tracks; A method having the following.
[0145] EEE29. A method for training a deep learning based system for automatic multi-track mixing based on a plurality of input audio tracks, the system having one or more instances of a first deep learning based network and one or more instances of a second deep learning based network, the method comprising: a training phase for training the second network, the training phase for training the second network comprising: obtaining at least one first training set as input, said first training set including, for each audio signal, at least one transformation parameter for signal processing of said audio signal and a respective predetermined processed audio signal; inputting the first training set into the second network; iteratively training the second network to predict each processed audio signal based on the audio signals and the transformation parameters in the first training set; and training the second network based on at least one first loss function indicative of a difference between the given processed audio signal and each of its predictions; method.
[0146] EEE30. The method further includes a training phase for training the first network, the training phase for training the first network comprising: obtaining at least one second training set as input, said second training set including a plurality of subsets of audio tracks and, for each subset, a predetermined mix of each of the audio tracks in said subset; inputting the second training set into the first network; iteratively training the first network to predict the mix of each of the subset of audio tracks in the second training set; and training the first network based on at least one second loss function indicative of the difference between a given mix of the audio tracks and their respective predictions; Method as described in EEE29.
[0147] EEE31. The training phase of training the first network begins after the training phase of training the second network ends. Method according to EEE30.
[0148] EEE32. A program having instructions which, when executed by a processor, cause the processor to perform the steps of the methods set forth in any one of EEE1 and EEE28-31.
[0149] A computer-readable storage medium storing a program described in EEE33.EEE32.
[0150] EEE34. An apparatus having a processor and a memory coupled to the processor, the processor being adapted to cause the apparatus to perform the steps of the method according to any one of EEE1 and EEE28-31, Device.
[0151] [Cross-reference to cross-application] This application claims priority from the following priority applications: Spanish (ES) Patent Application No. P202030604 (Reference No. D20041EP), filed June 22, 2020; U.S. Provisional Patent Application No. 63 / 072,762 (Reference No. D20041USP1#), filed August 31, 2020; U.S. Provisional Patent Application No. (Reference No. D20041USP2), filed October 15, 2020; and European Patent Application No. 20203276.9 (Reference No. D20041EP), filed October 22, 2020, which priority applications are incorporated herein by reference.
Claims
1. 1. A deep learning based system for automatic multi-track mixing based on multiple input audio tracks, comprising: one or more instances of a deep learning based controller network; one or more instances of a deep learning-based transformation network; and the controller network is configured to generate parameters used for the automatic multi-track mixing based on the input audio tracks; the transformation network is configured to apply signal processing and at least one mixing gain to the input audio tracks based on the parameters to generate an output mix of the input audio tracks; the controller network and the transformation network are trained separately, and the controller network is trained based on the pre-trained transformation network; system.
2. the output mix is a stereo mix, The system of claim 1 .
3. the controller network having a first stage and a second stage; Generating the parameters by the controller network includes: said first step comprising mapping each of said input audio tracks to a respective feature space representation; the second step generating parameters to be used by the transformation network based on the feature space representation; having 3. The system according to claim 1 or 2.
4. The second step of generating the parameters used by the transformation network includes: generating a combined representation based on the feature space representation of the input audio track; generating parameters for use by the transformation network based on the combined representation; having The system of claim 3.
5. generating the combined representation comprises an averaging process on the feature space representation of the input audio track; The system of claim 4.
6. the controller network is trained based on at least one loss function indicative of the difference between a given mix of audio tracks and their respective predictions; A system according to any one of claims 1 to 5.
7. The controller network obtaining at least one first training set as input, said first training set including a plurality of subsets of audio tracks and, for each subset, a predetermined mix of each of the audio tracks in said subset; inputting the first training set into the controller network; iteratively training the controller network to predict the mix of each of the subset of audio tracks in the first training set; Trained by the training is based on at least one first loss function indicative of the difference between a given mix of the audio tracks and their respective predictions; A system according to any one of claims 1 to 6.
8. the predicted mix of the audio track is a stereo mix, the first loss function is a stereo loss function and is configured to be invariant under left and right channel reassignment; The system of claim 7.
9. Training the controller network to predict the mix of the audio tracks includes, for each subset of audio tracks: generating, by the controller network, a plurality of prediction parameters according to the subset of audio tracks; providing the prediction parameters to the transformation network; generating, by said transformation network, a prediction of a mix of said subset of audio tracks based on said prediction parameters and said subset of audio tracks; having 9. A system according to claim 7 or 8.
10. the number of instances of said transformation network is equal to the number of said input audio tracks; the transformation network is configured to perform signal processing on each input audio track based at least in part on the parameters to generate a respective processed output; the processed output includes a left channel and a right channel; the output mix is generated based on the processed output. A system according to any one of claims 1 to 9.
11. The system further comprises a routing component; the routing component is configured to generate a plurality of bus level mixes based on the processing output; the output mix is generated based on the bus level mix; The system of claim 10.
12. the controller network is further configured to generate parameters for the routing component. The system of claim 11.
13. the one or more instances of the translation network are a first set of one or more instances of the translation network; the system further comprising a second set of one or more instances of the translation network; the number of instances in the second set of one or more instances of the transformation network is determined according to the number of bus level mixes.
13. A system according to claim 11 or 12.
14. the controller network is further configured to generate parameters for the second set of one or more instances of the transformation network. The system of claim 13.
15. the system is further configured to generate a left mastering mix and a right mastering mix based on the bus level mix; The system further comprises a pair of instances of the conversion network; the pair of instances of the transformation network are configured to generate the output mix based on the left mastering mix and the right mastering mix.
15. A system according to claim 13 or 14.
16. the controller network is further configured to generate parameters for the pair of instances of the transformation network. The system of claim 15.
17. The conversion network includes: - obtaining at least one second training set as input, said second training set comprising a plurality of audio signals and, for each audio signal, at least one transformation parameter for signal processing of said audio signal and a respective predetermined processed audio signal; inputting the second training set into the transformation network; iteratively training the transformation network to predict each processed audio signal based on the audio signal and the transformation parameters; Trained by the training is based on at least one second loss function indicative of a difference between the predetermined processed audio signal and each of its predictions; 17. A system according to any one of claims 1 to 16.
18. the parameters generated by the controller network are human and / or machine interpretable parameters; 18. A system according to any one of claims 1 to 17.
19. the parameters generated by the controller network include control parameters and / or panning parameters; 19. A system according to any one of claims 1 to 18.
20. the controller network and / or the transformation network comprises at least one neural network; The neural network has linear layers and / or multi-layer perceptrons (MLPs).
20. A system according to any one of claims 1 to 19.
21. The neural network is a convolutional neural network (CNN).
21. The system of claim 20.
22. the system having multiple instances of the controller network in a weight-sharing configuration and / or multiple instances of the transformation network in a weight-sharing configuration; 22. A system according to any one of claims 1 to 21.
23. 1. A method of operating a deep learning based system for automatic multi-track mixing based on a plurality of input audio tracks, the system having one or more instances of a deep learning based controller network and one or more instances of a deep learning based transformation network, the method comprising: generating, by the controller network, parameters to be used for the automatic multi-track mixing based on the input audio tracks; applying signal processing and at least one mixing gain to the input audio tracks based on the parameters by the transformation network to generate an output mix of the input audio tracks; and the controller network and the transformation network are trained separately, and the controller network is trained based on the pre-trained transformation network; method.
24. The transformation network is a trained transformation network that has been trained by a training phase, and the training phase used to train the transformation network is obtaining at least one first training set as input, said first training set comprising a plurality of audio signals and, for each audio signal, at least one transformation parameter for signal processing of said audio signal and a respective predetermined processed audio signal; inputting the first training set into the transformation network; iteratively training the transformation network to predict each processed audio signal based on the audio signals and the transformation parameters in the first training set; and training the transformation network based on at least one first loss function indicative of a difference between the predetermined processed audio signal and each of its predictions; 24. The method of claim 23.
25. The controller network is a trained controller network that has been trained by a training phase, the training phase used to train the controller network being: obtaining at least one second training set as input, said second training set including a plurality of subsets of audio tracks and, for each subset, a predetermined mix of each of the audio tracks in said subset; inputting the second training set into the controller network; iteratively training the controller network to predict the mix of each of the subset of audio tracks in the second training set; and the training of the controller network is based on at least one second loss function indicative of the difference between a predetermined mix of the audio tracks and their respective predictions; 25. The method of claim 24.
26. the training phase of training the controller network begins after the training phase of training the transformation network ends; 26. The method of claim 25.
27. A program comprising instructions which, when executed by a processor, cause the processor to carry out the steps of the method according to any one of claims 23 to 26.
28. A computer-readable storage medium storing the program according to claim 27.
29. 1. An apparatus having a processor and a memory coupled to the processor, The processor is adapted to execute a program stored in the memory to cause the device to perform the steps of the method of any one of claims 23 to 26. Device.
Citation Information
Patent Citations
Audio Mixer System
JP2016521925A
Improved convolutions of digital signals using a bit requirement optimization of a target digital signal
US20200321975A1
Automated music production
WO2019121574A1