Audio source separation processing pipeline system and method - Patents.com

JP2024540244A5Active Publication Date: 2025-11-05WINGNUT FILMS PROD LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024525939
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-23
Filing Date
2022-10-27
Publication Date
2025-11-05
Estimated Expiration
2042-10-27

AI Technical Summary

Technical Problem

Existing audio source separation techniques are not optimized to generate high-quality audio stems from low-quality, single-track audio recordings, particularly those with noisy mixtures from older sound recordings.

Method used

An improved audio source separation system and method that utilizes machine learning models to separate and enhance audio components, such as speech and instruments, by applying a processing recipe involving multiple source separation processes, neural networks to reduce artifacts, and a self-iterative training process to refine the models.

Benefits of technology

The system effectively separates and enhances audio sources from low-quality, single-track recordings, reducing artifacts like clicks, harmonic distortion, and broadband noise, resulting in high-fidelity audio stems suitable for modern sound systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A system and method for audio source separation includes receiving a single-track audio input sample having an unknown mixture of audio signals generated from multiple audio sources, and separating one or more of the audio sources from the single-track audio input sample using a sequential audio source separation model. Separating one or more of the audio sources may include defining a processing recipe including a multiple source separation process configured to receive the audio input mixture and output one or more separated source signals and a remaining complementary signal mixture, and processing the single-track audio input sample according to the processing recipe to generate a plurality of audio stems separated from the unknown mixture of audio signals.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This disclosure claims priority to U.S. Patent Application No. 17 / 848,341, filed June 23, 2022, and claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 272,650, filed October 27, 2021, both of which are incorporated by reference herein in their entireties.

[0002] The present disclosure relates generally to systems and methods for audio source separation, and more particularly, to systems and methods for separating and enhancing an audio source signal from an audio mixture, such as a single track audio mixture. [Background technology]

[0003] Audio mixing is the process of combining multiple audio recordings to produce an optimized mixture for playback in one or more desired sound formats, such as mono, stereo, or surround sound. In applications requiring high quality sound production, such as sound production for music and movies, audio mixtures are typically produced by mixing separate high quality recordings. These separate recordings are often generated in a controlled environment, such as a recording studio with optimized sound effects and high quality recording equipment.

[0004] Often, some of the source audio may be of poor quality and / or contain a mixture of the desired audio source and unwanted noise. In modern audio post-production, it is common to re-record audio when the original recording lacks the desired quality. For example, in music recording, vocal or instrument tracks may be recorded and mixed with the previous recording. In sound post-production for movies, it is common to bring actors back into the studio to re-record their lines and add other audio (e.g., sound effects, music) to the mix.

[0005] However, in some applications, it is desirable to faithfully convert the original audio source into a high-quality audio mix. For example, movies, music, television broadcasts, and other audio recordings can date back more than 100 years. The source audio was recorded on older, lower-quality equipment and may contain a low-quality mixture of desired audio and noise. For many recordings, a single track / mono audio mix is ​​the only audio source available to produce an optimized mix for playback on modern sound systems.

[0006] One approach to processing an audio mixture is to separate the audio mixture into a set of separate audio source components and generate a separate audio stem for each component of the audio mixture. For example, a music recording may be separated into a vocal component, a guitar component, a bass component, and a drum component. Each of the separate components may then be enhanced and mixed to optimize playback.

[0007] However, existing audio source separation techniques are not optimized to generate the high quality audio stems required to produce high fidelity output for the music and film industries. Audio source separation is particularly difficult when the audio sources are low quality from older sound recordings and include single track noisy sound mixtures.

[0008] In view of the foregoing, there is a continuing need for improved audio source separation systems and methods, particularly for the generation of high fidelity audio from lower quality audio sources.

[0009] It is an object of at least the preferred embodiments to address at least some of the disadvantages discussed above. An additional or alternative object is to at least provide the public with a useful alternative to conventional techniques. Summary of the Invention [Means for solving the problem]

[0010] Improved audio source separation systems and methods are disclosed herein. In various implementations, a single track audio recording is provided to an audio source separation system configured to separate and separate various audio components, such as speech and individual instruments, into high fidelity stems (e.g., discrete or grouped collections of audio sources mixed together).

[0011] In some implementations, an audio source separation system includes a first machine learning model that is trained to separate a single-track audio recording into stems, including speech, complements, and artifacts such as "clicks." Additional machine learning models may then be used to refine the speech stems by removing processing artifacts from the speech and / or fine-tuning the first machine learning model.

[0012] The term "comprising" as used herein means "consisting at least in part of." When interpreting each phrase in this specification containing the term "comprising," features other than the one or more prefaced by the term may also be present. Related terms such as "comprise" and "comprises" are to be interpreted in the same manner.

[0013] In various implementations, the method includes receiving a single track audio input sample comprising an unknown mixture of audio signals generated from a plurality of audio sources; defining a processing recipe comprising a multiple source separation process configured to receive the audio input mixture and output one or more separated source signals and a remaining complementary signal mixture; and processing the single track audio input sample according to the processing recipe to generate a plurality of audio stems separated from the unknown mixture of audio signals, using a sequential audio source separation model to separate one or more of the audio sources from the single track audio input sample.

[0014] The method may further define a processing recipe by processing the single track audio input samples using a first processing order of the multiple source separation processes for separating the audio stems to generate a first set of audio stems, processing the single track audio input samples using a second processing order of the multiple source separation processes different from the first processing order for separating the audio stems to generate a second set of audio stems, and evaluating the first set of audio stems and the second set of audio stems to determine which of the first processing order and the second processing order should be included in the processing recipe.

[0015] The processing recipe of the method may further include post-processing the multiple audio stems to remove and / or mitigate artifacts introduced by the sequential audio source separation model, where the post-processing includes processing the multiple audio stems through one or more neural networks trained to remove and / or mitigate artifacts comprising clicks, harmonic distortion, ghosting, and / or broadband noise.

[0016] The processing recipe of the method may further include executing multiple source separation processes in a sequential, branching processing order, where a first source separation process is configured to receive a single track audio input sample and output one or more source separation signals and a remaining complementary signal mixture, and each subsequent separation process is configured to receive an output signal from a prior source separation process, and each node of the processing recipe is configured to separate a particular source class from the input audio mixture.

[0017] The processing recipe of the method may further include a step of post-processing the multiple audio stems by combining two or more of the multiple audio stems having a common source class generated by different source separation processes in the processing recipe, and / or a step of processing one or more of the multiple audio stems and mitigating artifacts and / or noise introduced by the sequential audio source separation model by applying a neural network model trained to clean up artifacts on the audio stems separated according to the processing recipe.

[0018] At least one source separation process of the method may further include separating and outputting the mixture of the speech source and a residual complement comprising music and noise.

[0019] The method may further include evaluating output audio stems from a plurality of sequential branching processing orders and assessing an optimized processing order, the evaluating step including processing the input mixture, separating a first source class and a first residual complement, then processing the first residual complement, separating a second source class, and generating a first set of audio output stems, the evaluating step further including processing the input mixture, separating the second source class and a second residual complement, then processing the second residual complement, separating the first source class, and generating a second set of audio output stems. The method may further include comparing the first set of audio output stems and the second set of audio output stems and determining a processing order that generates a higher quality audio output stem.

[0020] The step of separating one or more of the audio sources from the single track audio input sample using the sequential audio source separation model of the method may further include a user-guided and / or self-iterative process configured to progressively separate the sources from the unknown mixture of audio signals by selecting audio enhancements and / or audio enhancement parameters and fine-tuning the sequential audio source separation model to enable matching of source signals in the unknown mixture of audio signals, and identifying source classes in the unknown mixture of audio signals and identifying a general separation model for use in a processing recipe.

[0021] The processing recipe of the recipe may further comprise a set of source class and output stems and corresponding source separation models arranged in a hierarchical branching sequence, where at least one class of sources is separated into multiple source class separation stems and / or the processing recipe includes a step of generating a mixture of multiple source class separation stems as a source class output.

[0022] The method may further include training a sequential audio source separation model using the single track audio input sample by providing the single track audio input sample to a generic source separation model, generating a plurality of initial audio stems corresponding to one or more of the plurality of audio sources, and retraining the generic source separation model, at least in part, using one or more of the plurality of initial audio stems, the generic source separation model comprising a plurality of neural network models, each neural network model configured to receive a single channel audio input sample and output one or more source-separated audio stems comprising a source class and a remaining complementary signal mixture.

[0023] The training of the sequential audio source separation model of the method may further include training a plurality of neural network models in an order defined by a processing recipe, and / or the processing recipe comprises a hierarchical branching sequence, with each branch comprising one or more of the plurality of neural network models. The training of the sequential audio source separation model of the method may further include re-iteratively training the sequential audio source separation model based, at least in part, on a plurality of audio stems generated from a prior iteration of the sequential audio source separation model, determining a metric associated with the source-separated audio stem and evaluating one or more of the plurality of audio stems by comparing the metric to one or more threshold parameters, and / or adding the one or more source-separated audio stems to a training dataset to train the sequential audio source separation model based on the evaluating step.

[0024] The training step may further include generating a training data set of artifacts generated by the process recipe to train the one or more neural networks, receiving a source separation output with the artifacts, and generating an enhanced output in which the one or more artifacts have been reduced and / or removed according to the process recipe.

[0025] In various implementations, the system includes a memory component that stores machine-readable instructions and a logic device configured to execute the machine-executable instructions to separate one or more audio sources from a single track audio input sample comprising an unknown mixture of audio signals generated from the multiple audio sources using a sequential audio source separation model by: defining a processing recipe including a multiple source separation process configured to receive an audio input mixture and output one or more separated source signals and a remaining complementary signal mixture; and processing the single track audio input sample according to the processing recipe to generate multiple audio stems separated from the unknown mixture of audio signals.

[0026] The logic device may be further configured to define a processing recipe by: processing the single track audio input samples using a first processing order of the multiple source separation processes for separating the audio stems to generate a first set of audio stems; processing the single track audio input samples using a second processing order of the multiple source separation processes different from the first processing order for separating the audio stems to generate a second set of audio stems; and evaluating the first set of audio stems and the second set of audio stems to determine which of the first processing order and the second processing order should be included in the processing recipe.

[0027] The logic device may further be configured to define a processing recipe by post-processing the multiple audio stems to remove and / or mitigate artifacts introduced by the sequential audio source separation model, where the post-processing includes processing the multiple audio stems through one or more neural networks trained to remove and / or mitigate artifacts comprising clicks, harmonic distortion, ghosting, and / or broadband noise.

[0028] The system may be further defined in that at least one source separation process includes separating and outputting a mixture of a speech source and a residual complement comprising music and noise.

[0029] The logic device of the system may be further configured to define the processing by executing multiple source separation processes in a sequential, branching processing order, where a first source separation process is configured to receive a single track audio input sample and output one or more source separation signals and a remaining complementary signal mixture, and where each subsequent separation process is configured to receive an output signal from a prior source separation process, and where each node of the processing recipe is configured to separate a particular source class from the input audio mixture.

[0030] The logic device may be further configured to evaluate the output audio stems from the multiple sequential divergent processing orders and assess an optimized processing order by: processing the input mixture and separating the first source class and the first residual complement, then processing the first residual complement, separating the second source class, and generating a first set of audio output stems; processing the input mixture and separating the second source class and the second residual complement, then processing the second residual complement, separating the first source class, and generating a second set of audio output stems; and comparing the first set of audio output stems and the second set of audio output stems and determining a processing order that generates a higher quality audio output stem.

[0031] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter or to limit the scope of the claimed subject matter. A more extensive presentation of the features, details, utilities, and advantages of the present methods as defined in the claims is provided in the following written description of various implementations of the disclosure and illustrated in the accompanying drawings.

[0032] Where patent specifications, other external documents, or other sources of information are referenced herein, this is generally for the purpose of providing a context for discussing the features of the present invention. Unless specifically stated otherwise, reference to such external documents or such sources of information is not to be construed as an admission that such documents or such sources are prior art or form part of the common general knowledge in the art in any jurisdiction. [Brief description of the drawings]

[0033] Aspects of the present disclosure and their advantages can be better understood by reference to the following drawings and the detailed description that follows. It should be understood that like reference numbers are used to identify like elements illustrated in one or more of the figures, and that the illustrations therein are intended to illustrate implementations of the present disclosure, and not to limit it. The components in the drawings are not necessarily to scale, emphasis instead being placed upon clearly illustrating the principles of the present disclosure.

[0034] [Figure 1] FIG. 1 illustrates an audio source separation system and process, according to one or more implementations.

[0035] [Diagram 2] FIG. 2 illustrates relevant elements of the system and process of FIG. 1, in accordance with one or more implementations.

[0036] [Diagram 3] FIG. 3 is a diagram illustrating a machine learning dataset and training data loader, according to one or more implementations.

[0037] [Figure 4] FIG. 4 illustrates an example machine learning training system that includes a self-iterating dataset generation loop, according to one or more implementations.

[0038] [Diagram 5] FIG. 5 illustrates an example operation of a data loader for use in training a machine learning system, according to one or more implementations.

[0039] [Figure 6] FIG. 6 illustrates an exemplary machine learning training method, according to one or more implementations.

[0040] [Figure 7]FIG. 7 illustrates an example machine learning training method, including a training mixture example, according to one or more implementations.

[0041] [Figure 8] FIG. 8 illustrates an example machine learning process, according to one or more implementations.

[0042] [Figure 9] FIG. 9 illustrates an example post-processing model configured to clean up artifacts introduced by machine learning processes, according to one or more implementations.

[0043] [Figure 10] FIG. 10 illustrates an exemplary user-guided self-iterative training loop, according to one or more implementations.

[0044] [Figure 11A] FIG. 11 comprises FIGS. 11A and 11B, which illustrate an example machine learning application, according to one or more implementations. [Figure 11B] FIG. 11 comprises FIGS. 11A and 11B, which illustrate an example machine learning application, according to one or more implementations.

[0045] [Figure 12] FIG. 12 illustrates an example machine learning processing application, according to one or more implementations.

[0046] [Figure 13A] FIG. 13 comprises FIGS. 13A, 13B, 13C, 13D, and 13E illustrating one or more examples of a multi-model recipe processing system according to one or more implementations. [Figure 13B] FIG. 13 comprises FIGS. 13A, 13B, 13C, 13D, and 13E illustrating one or more examples of a multi-model recipe processing system according to one or more implementations. [Figure 13C] FIG. 13 comprises FIGS. 13A, 13B, 13C, 13D, and 13E illustrating one or more examples of a multi-model recipe processing system according to one or more implementations. [Figure 13D] FIG. 13 comprises FIGS. 13A, 13B, 13C, 13D, and 13E illustrating one or more examples of a multi-model recipe processing system according to one or more implementations. [Figure 13E] FIG. 13 comprises FIGS. 13A, 13B, 13C, 13D, and 13E illustrating one or more examples of a multi-model recipe processing system according to one or more implementations.

[0047] [Figure 14] FIG. 14 illustrates an exemplary audio processing system according to one or more implementations.

[0048] [Figure 15] FIG. 15 illustrates an example neural network that may be used in one or more of the implementations of FIGS. 1-14 according to one or more implementations. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0049] Detailed Description In the following description, various implementations will be described. For purposes of explanation, specific configurations and details are set forth to provide a thorough understanding of the implementations. However, it will also be apparent to those skilled in the art that the implementations may be practiced without the specific details. Furthermore, well-known features may be omitted or simplified to avoid obscuring the implementations being described.

[0050] Improved audio source separation systems and methods are disclosed herein. In various implementations, a single-track (e.g., undifferentiated) audio recording is provided to an audio source separation system configured to separate and separate various audio components, such as speech and instruments, into high-fidelity stems, i.e., discrete or grouped collections of audio sources mixed together. In various implementations, the single-track audio recording contains an unidentified audio mixture (e.g., the audio sources, recording environment, and / or other aspects of the audio mixture are unknown to the audio source separation system), and the audio source separation system and method are adapted to identify and / or separate audio sources from the unidentified audio mixture in a self-iterative training and fine-tuning process.

[0051] The systems and methods disclosed herein may be implemented on at least one computer-readable medium carrying instructions that, when executed by at least one processor, cause the at least one processor to perform any of the method steps disclosed herein. Some implementations relate to a computer system including at least one processor and a memory that stores instructions that, when executed by the at least one processor, cause the at least one processor to perform any of the method steps disclosed herein. In various implementations, the models described herein may be implemented as stored data and software modules and / or code that operates on the stored data.

[0052] In some cases, training a model involves processing a data structure or structures to form a new data structure that a component of a computer system can access and use as a model. For example, an artificial intelligence system may include a computer with one or more processors, a program code memory, a writable data memory, and several inputs / outputs. The writable data memory may hold several data structures corresponding to trained or untrained models. Such data structures may represent one or more layers of nodes of a neural network and links between nodes of different layers and weights for at least some of the links between nodes. In other cases, different types of data structures may represent the model.

[0053] In some cases, when referring to training a model, feeding a model, and / or having a model take in inputs and provide outputs, this may refer to the action of a computer capable of reading a writable data memory that contains the model and executes program code for working with the model. For example, the model may be trained with a set of training data, which may be the training examples themselves and / or the training examples and corresponding ground truth. Once trained, the model may be usable to make decisions about the examples provided to the model. This may be done by the computer receiving, at the input, input data representing the examples, performing a process using the examples and the model, and outputting, at the output, output data representing and / or indicating a decision made by or based on the model.

[0054] In a very specific example, an artificial intelligence system may have a processor that reads in a number of photos of cars and reads in ground truth data indicating "these are cars". The processor may read in a number of photos of street lamps and the like and read in ground truth data indicating "these are not cars". In some cases, the model is trained on the input data itself, without being provided with ground truth data. The result of such processing may be a trained model. The artificial intelligence system may then, with the model trained, be provided with an image without any indication of whether this is an image of a car, and may output an indication of a determination of whether this is an image of a car.

[0055] To process audio signals, data, recordings, etc., the input may or may not be the audio data itself, and some ground truth data about the audio data. Then, once trained, the artificial intelligence system may receive some unknown audio data and output decision data about the unknown audio data. For example, the output decision data may relate to extracted notes, stems, frequencies, etc., or other decisions or AI-decision observations of the input audio data.

[0056] The resulting data structure corresponding to the trained model can then be ported or distributed to other computer systems, which can then use the trained model. Once trained, the computer code in the program memory, when executed by a processor, can receive an image at the input, and based on the fact that the data structure represents the trained AI model being trained, the program code can process the input and output a decision regarding the nature of the input. In some implementations, an AI model may comprise a data structure and program code that are intertwined and not easily separated.

[0057] The model may be represented by a set of weights assigned to edges in the neural network graph, program code and data including instructions on the neural network graph and how to interact with the graph, mathematical expressions such as regression or classification models, and / or other data structures as may be known in the art. A neural model (or neural network, or neural network model) may be embodied as a data structure that shows or represents a set of connected nodes, often referred to as neurons, many of which may be data structures that mimic or simulate the signal processing performed by biological neurons. Training may include updating parameters associated with each neuron, such as weights and / or functions of other neurons to which it is connected and the neuron's input and output. In practice, a neural model may pass and process input data variables in some way to generate output variables to achieve some goal, for example to generate a binary classification of whether an input image or input data set fits into a certain category. The training process may involve complex calculations (e.g., calculating gradient updates and then using the gradients to update parameters layer by layer). The training may be performed using some parallel processing.

[0058] Any of the models for audio source separation described herein may also be referred to, at least in part, by a number of terms including, for example, an “audio source separation model,” “recurrent neural network,” “RNN,” “deep neural network,” “DNN,” “inference model,” or “neural network.”

[0059] 1 and 2 illustrate an audio source separation system and process according to one or more implementations of the present disclosure. The audio processing system 100 includes a core machine learning system 110, a core model operation 130, and a modified recurrent neural network (RNN) class model 160. In the illustrated implementation, the core machine learning system 110 implements an RNN class audio source separation model (RNN-CASSM) 112, depicted in a simplified representation in FIG. 1. As shown, the RNN-CASSM 112 receives a signal input 114, which is input to a time domain encoder 116. The signal input 114 includes a single channel audio mixture, which may be received from a stored audio file accessed by the RNN-CASSM network, an audio input stream received from a separate system component, or other audio data source. The time domain encoder 116 models the input audio signal in the time domain and estimates audio mixture weights. In some implementations, the time domain encoder 116 segments the audio signal input into separate waveform segments that are normalized for input to a 1D convolutional coder. The RNN class mask network 118 is configured to estimate a source mask for separating the audio sources from the audio input mix. The mask is applied to the audio segments by a source separation component 120. The time domain decoder 122 is configured to reconstruct the audio sources, which are then available for output through a signal output 124.

[0060] In an exemplary implementation, the RNN-CASSM 112 is modified for operation with the core model operations 130, which will be described herein. Referring to block 1.1, the audio source data is sampled using a 48 kHz sample rate in various implementations. Thus, the RNN-CASSM 112 may be trained at higher than audible frequencies to recognize separate stems of audio in the lower frequency range. It has been observed that implementations of speech separation models trained at sample rates lower than 48 kHz produce lower quality audio separation in audio samples recorded with older equipment (e.g., Nagra™ equipment from the 1960s). Various steps disclosed herein, such as training the source separation model, setting appropriate hyperparameters for that sample rate, and operating the signal processing pipeline at the 48 kHz sample rate, are performed at the 48 kHz sample rate. Other sampling rates, including oversampling, may also be used in other implementations, consistent with the teachings of this disclosure. In block 1.2, the encoder / decoder framework (eg, the time domain encoder 116 and the time domain decoder 122) is set to a step size of one sample (eg, the input signal sample rate).

[0061] Referring to block 1.3, a modified RNN-CASSM 160 is generated that extends beyond the source separation and noise reduction of conventional implementations. It is observed that processing audio mixtures with trained RNN-CASSM can sometimes result in undesirable click artifacts, harmonic artifacts, and broadband noise artifacts. To address this, the audio processing system 100 applies modifications to the RNN-CASSM network 112, as mentioned in blocks 1.3a, 1.3b, and / or 1.3c, to reduce and / or avoid the artifacts and provide other advantages. The same modifications may also be used for more transformative types of learned processing as well.

[0062] The model operation associated with the modified RNN-CASSM 160 will now be described in more detail with reference to blocks 1.3ac. In block 1.3a, at least some of the encoder and decoder layers are removed and not used in the modified RNN-CASSM. It is observed that these removed layers may be redundant in the trained filter when the step size is one sample. In block 1.3b, the step of applying a mask (e.g., component 120) is also removed for at least some part of the audio processing. Thus, the modified RNN-CASSM 160 may perform audio source separation without applying a mask. The masking step potentially contributes to click artifacts that are often present in the generated audio stem. To address this, the audio processing system may omit some or all of the masking and instead use the output of the RNN class mask network 118 in a more direct manner (e.g., training the RNN to output one or more separated audio sources). In block 1.3c, a window function is applied to the overlap-add step of an RNN class mask network (e.g., RNN network 162). When the model output has linear harmonic series "banding" artifacts at frequencies related to the overlap-add function segment length, the audio processing system can expand the window function across each overlapping segment to smooth out hard edges when reconstructing the audio signal.

[0063] Referring to block 1.4, another core model operation 130 is the use of a separation strength parameter, which allows control over the strength of the separation mask that is applied to the input signal to generate the separated sources. To provide direct control over the strength of the separation mask that is applied to the input signal to generate the separated sources, a parameter is introduced during the forward pass of the model that determines how strongly the separation mask is applied. In one embodiment, the separation strength parameter is f(M)=M Setc., where the mask M has values ​​[0, l] and s is a separation strength parameter. In this example, values ​​of s>1.0 result in lower mask values ​​and fewer target sources when the mask is applied to the input mixture, and values ​​of s<1.0 result in higher mask values ​​and a combination of the target sources and complementary and noise components in the signal.

[0064] The separation strength parameter may be implemented as a helper function of an automated version of the Self-Iterative Process Training (SIPT) algorithm, which will be described in more detail below. It should be understood that the RNN-CASSM 112 may implement one or more of the core model operations 130 disclosed herein and may include additional operations consistent with the teachings of this disclosure.

[0065] With reference to Figures 3-5, an exemplary process for training an RNN-CASSM network for audio source separation will now be described. The training process includes a plurality of labeled machine learning training datasets 310 configured to train a network for audio source separation as described herein. For example, in some implementations, the training dataset may include an audio mix and ground truth labels that identify source classes to be separated from the audio mix. In other implementations, the training dataset may include separated audio stems having audio artifacts (e.g., clicks) generated by the source separation process and / or one or more audio enhancements (e.g., reverb, filters), and ground truth labels that identify enhanced audio stems in which the identified audio artifacts and / or enhancements have been removed.

[0066] In operation, the network is trained by feeding labeled audio samples into the network. In various implementations, the network includes multiple neural network models that can be trained separately for specific source separation tasks (e.g., separating speech, separating foreground speech, separating drums, removing artifacts, etc.). Training involves a forward pass through the network to generate audio source separation data. Each audio sample is labeled with a "ground truth" that defines the expected output, which is compared to the generated audio source separation data. If the network mislabels an input audio sample, a backward pass through the network may be used to adjust the parameters of the network to correct for the misclassification. In various implementations, [ka] The output estimate, denoted as Y, is compared to the ground truth, denoted as Y, using a regression loss, such as an L1 loss function (e.g., least absolute deviation), an L2 loss function (e.g., least squares), a scale-invariant signal-to-distortion ratio (SISDR), a scale-dependent signal-to-distortion ratio (SDSDR), and / or other loss functions as known in the art. After the network is trained, a validation dataset (e.g., a set of labeled audio samples not used in the training process) may then be used to measure the accuracy of the trained network. The trained RNN-CASSM network may then be implemented in a runtime environment to generate separate audio source signals from the audio input stream. The generated separate audio source signals may also be referred to as generated multiple audio stems. The generated multiple audio stems may correspond to one or more audio sources of the multiple audio sources of the audio input stream.

[0067] In various implementations, the modified RNN-CASSM network 160 is trained based on multiple (e.g., thousands) of audio samples, including audio samples representing multiple speakers, instruments, and other audio source information under various conditions (including, for example, various noisy conditions and audio mixes). From the error between the separated audio source signals and the labeled ground truth associated with the audio samples from the training dataset, the deep learning model learns parameters that enable the model to separate the audio source signals. The RNN-CASM 120 and the modified RNN-CASM 160 may also be referred to as the inference model 120 and the modified inference model 160 and / or the trained audio separation model 120 and the updated audio separation model 160.

[0068] 3 is a schematic diagram of a machine learning dataset 310 and a training data loader 350 as may be relevant for the dataset and dataset manipulation during training by an audio processor that may provide improved and / or useful functionality in output signal quality source separation. In the illustrated implementation, an RNN-CASSM network is trained to separate audio sources from a single track recording of a music recording session that includes a mixture of people singing / talking, instruments being played, and various environmental noises.

[0069] Referring to block 2.1, the training data set used in the illustrated implementation includes a 48 kHz speech data set. For example, the 48 kHz speech data set may include the same speech recorded simultaneously at various microphone distances (e.g., a close microphone and a more distant microphone). In one test implementation, 85 different speakers were included in the 48 kHz speech data set with more than 20 minutes of speech per speaker. In various implementations, the exemplary speech data set may be created using a large number of speakers, such as 10, 50, 85, or more, from adult male and female speakers, with extended speech periods, such as 10 minutes, 20 minutes, or more, and recordings at 48 kHz or higher sampling rates. It should be understood that other data set parameters may also be used to generate a training data set in accordance with the teachings of the present disclosure.

[0070] Referring to block 2.2, the training dataset further includes a non-speech music and noise dataset that includes segments of input audio, such as digitized mono audio recordings originally recorded on analog media. In some implementations, the dataset may include segments of recorded music, non-vocal sounds, background noise, audio media artifacts from digitized audio recordings, and other audio data. Using the dataset, an audio processing system can more easily separate the voice of a speaker of interest from other voices, music, and background noise in the digitized recording. In some implementations, this may include using manually collected segments of the recordings that are manually annotated as lacking speech or another audio source class, and labeling those segments accordingly.

[0071] Referring to block 2.3, a dataset is generated and refined using a progressive self-iterative dataset generation process using a target unidentified mixture (e.g., a mixture unknown to the audio processing system). The generated dataset may include a labeled dataset generated by processing an unlabeled dataset (e.g., a target unidentified mixture to be source separated) through one or more neural network models to generate an initial classification. This "coarsely separated data" is then processed through a cleanup step configured to select among the "coarsely separated data" to keep the most useful "coarsely separated data" based on a usefulness metric. For example, the performance of the training dataset may be measured by applying a validation dataset to a model trained using various training datasets including the "coarsely separated data" and determining, based on a calculated validation error, data samples that contribute to better and poor performance. The usefulness metric may be implemented as a function that estimates a quality metric in the coarsely separated data to identify "low quality" fine-tuning data that should be discarded before the next fine-tuning iteration. For example, a moving root-mean-square (RMS) window function may be calculated on the output of the network to identify segments of the output where the RMS metric lies above a calibrated threshold over a certain (or minimum) duration of samples. This metric can be used, for example, to identify low amplitude segments of coarsely separated source outputs where artifacts are more likely to occur. The threshold and minimum duration may be user adjustable parameters, allowing tuning of the data to be discarded.

[0072] The progressively self-iterating dataset generated using the target unidentified mixture may be generated using a self-iterating dataset generation loop 420 as illustrated in FIG. 4. In various implementations, previous recordings containing the sources to be separated that are still unidentified by the trained source separation model are less likely to be successfully separated by the model. The existing dataset may not be substantial enough to train a robust separation model for the targeted sources, and the opportunity to capture new recordings of the sources may not exist. Instead of capturing new recordings of the sources, additional training data may be manually labeled from isolated instances of the sources in previous recordings. For example, isolated speech from an identified speaker, isolated audio of an identified instrument recorded with similar equipment in a similar environment, and / or other available audio segments that approximate the source being separated may be added to the training data manually and / or automatically (e.g., based on metadata labeling of the audio source, such as source identification, source class, and / or environment). This additional training data may be used to help fine-tune training the model to improve processing performance on previous recordings. However, this labeling process can involve a significant amount of time and manual effort, and there may not be enough isolated instances of a source in previous recordings to provide a sufficient amount of additional training data.

[0073] The illustrated generative refinement tool can overcome these difficulties. In one method, a coarse generic model 410 is trained on a generic training dataset. The generic training dataset may comprise labeled source audio data and labeled noise audio data. The generic model 410 may be referred to as a generic source separation model 410 or a trained audio source separation model 410. The training dataset may comprise a plurality of datasets, each of which may comprise labeled audio samples configured to train the system to address the source separation problem. The plurality of datasets may comprise a speech training dataset comprising a plurality of labeled speech samples and / or a non-speech training dataset comprising a plurality of labeled music and / or noise data samples. The available prior recordings containing the unidentified audio mixture to be separated are then processed with the generic model in process 422, resulting in two labeled datasets of isolated audio (e.g., audio stems), namely, a coarse separated prior recording source dataset 424 and a coarse separated prior recording noise dataset 426.

[0074] In various implementations, other training data sets may also be used that provide a set of labeled audio samples that are selected to train the system to solve a particular problem (e.g., speech vs. non-speech). In some implementations, for example, the trained data sets may include (i) music vs. sound effects vs. foley, (ii) data sets for various instruments in a band, (iii) multiple human speakers separated from each other, (iv) sources from room reverberation, and / or (v) other training data sets. The results are then culled (process 428) using a threshold metric to remove audio windows that are below a selected root-mean-square (RMS) level, which may be user-selectable. In some implementations, a running RMS may be calculated by segmenting the audio data into overlapping windows of equal duration and calculating the RMS for each window. The RMS level may be referred to as a usability metric or a quality metric, and the RMS level may be one option of alternative usability or quality metrics. The quality metric may be calculated based on multiple associated audio stems.

[0075] A new model is then trained in process 430 using the culled self-repetitive dataset (e.g., the culled result added to the audio training dataset, the culled self-repetitive dataset is also referred to as a culled dynamically evolving dataset) to train and generate an improved model 432 and improve its performance when processing the recording. The improved model 432 may be configured to reprocess the audio input stream and generate multiple improved audio stems. This improved model is an update of the trained audio source separation model. This improved model may be referred to as an updated audio source separation model 432.

[0076] In some implementations, the audio training dataset may be curated during an iterative fine-tuning process to remove data that is not relevant to the target input mixture. For example, the input mixture may identify / classify various sources within the input mixture, leaving certain other source categories that are not identified / classified and / or are otherwise not relevant to the source separation task. Training data associated with these "unrelevant" source categories (e.g., categories not found in the target mixture, categories identified by a user as not relevant to the source separation task, and / or other unrelevant source categories as defined by other criteria) may be culled from the audio training dataset, allowing the training dataset to become increasingly specific to the content of the target input mixture.

[0077] This process 420 is repeated iteratively, each of which can improve the separation quality of the model (e.g., fine-tuning to improve the accuracy and / or quality of source separation). In response to subsequent re-iterations, the process 420 may use additional RMS levels, whereby the additional RMS levels exceed the previous RMS levels. This process allows for more automated refinement of the initial general source isolation or separation model. It has been observed that looping at various stages shows improvement over larger related mixtures. The general model 410 and the improved model 432 may also be referred to as the inference model 410 and the modified inference model 432 and / or the trained audio separation model 410 and the updated audio separation model 432.

[0078] Improvements in separation quality (e.g., audio fidelity) can be measured by the system and / or assessed by a user providing feedback to the system through a user interface and / or overseeing one or more of the steps in process 420. In some implementations, process 420 may use a combination of algorithms and / or user ratings to calculate a Mean Opinion Score (MOS) to estimate separation quality. For example, an algorithm can estimate the amount of artifacting generated during the source separation operation, which in turn relates to the overall quality of the network's output. In some implementations, estimating the amount of artifacting generated during the source separation operation includes feeding the separated sources through a neural networking model trained to separate audio artifacts from the signal, allowing for a measurement of the presence and / or strength of such audio artifacts. The strength of the audio artifacts can be determined at each iteration and tracked between iterations to fine-tune the model. In some implementations, the iterative process continues until the estimated separation quality across iterations no longer improves and / or the estimated separation quality meets one or more predetermined quality thresholds.

[0079] Referring to blocks 2.4 and 2.4a, the machine learning training data loader 350 is configured to match the sound qualities (e.g., perceived distance from the microphone, filtering, reverberation, echo, nonlinear distortion, spectral distribution, and / or other measurable audio qualities) of the target unidentified mixture during training. The challenge with training an effective supervised source separation model is that the dataset example should ideally be curated to match as closely as possible the quality of the sources in the target mixture. For example, if a speaker should be isolated from the mixture in which they are speaking into a microphone in a reverberant hall, and the recording is captured from some distance with a recording device in the audience, a comparison may be made with how the same speaker might sound when recorded speaking directly into a microphone in a neutral, non-reverberant space, such as where a high-quality speaker dataset might have been recorded. The goal would then be to add enhancements to the high-quality speech dataset samples during training that generally result in lower deviations from the target input mixture. In this example, a "reverberation" extension may be added to simulate that of a hall, a "non-linear distortion" extension to simulate the voice of a speaker amplified by a sound system, and a "filter" extension to simulate the distance of the sound system from the recording device.

[0080] In various implementations, the solution includes a hierarchical mix bus scheme, including pre- and post-mixture extension modules. Creating an ideal target mixture would be tedious or impractical to create manually each time a new type of mixture is required for training. The hierarchical mix bus scheme allows easy definition of arbitrarily complex randomized "sources" and "noise" during supervised source separation training. The data loader approximately matches qualities such as filtering, reverberation, relative signal level training, and other audio qualities for improved source separation enhancement results. The machine learning data loader uses a hierarchical scheme that allows easy definition of "source" and "noise" mixtures that are dynamically generated from source data while training the model. The mix bus allows optional extensions such as reverb or filters with accompanying randomization parameters. Using appropriately classified dataset media as raw material, this allows easy setup of a training dataset that mimics the desired source separation target mixture.

[0081] An exemplary simplified schema representation 550 is illustrated in Figure 5. The training mixture schema includes separate options for sources and noise, including criteria such as dB ranges, probabilities associated with source determination, room impulse responses, filters, and other criteria.

[0082] Referring to block 2.4b, the data loader further provides pre / post mixture augmentations, including filters, nonlinear functions, convolutions, that are applied to the target mixture during training. In various implementations, relevant augmentations are identified and added to the training dataset, and the separated audio stems are post-processed (e.g., using the pipeline described herein). The source separation model can be trained with an additive goal of transforming the separated sources using the augmentations. In some implementations, the system can be trained to strictly isolate the source as it can be heard in the input mixture. In some implementations, the system may be further trained to improve the quality of some of the separated sources by applying appropriate augmentations, e.g., an augmentation that results in the least deviation from the target mixture when using the available training dataset. The deviations and appropriate augmentations may be estimated algorithmically and / or by user evaluation. For example, a voice in the input mixture may be filtered and difficult to understand due to being recorded behind a closed door. In this example, the separated sources may be augmented (e.g., to degrade the input audio dataset and generally match the sound of a voice behind a closed door). However, the target separated source outputs during training are not augmented in this embodiment (e.g., augmented speech inputs during training vs. corresponding high quality speech target output datasets), and thus the network is trained to approximate this same transformation by post-processing augmentation.

[0083] In some implementations, the target source may be a transformed version of its current representation in the mixture. For example, there may be a need to restore a bandwidth-limited recording to a more complete frequency spectrum, or to isolate and increase the close fidelity of an obscured background speaker. These needs may be user-determined, or left to the transformative model itself to resolve automatically based on input deviations from the target output training set. For example, if the model has been trained to output high-quality close-neighbor speech without much reverberation when using a randomly expanded speech dataset during training, inputting mixtures containing such high-quality close-neighbor speech without much reverberation may tend to result in minimal changes to those inputs. However, inputting mixtures containing speech that deviates from these qualities may tend to transform those input speech mixtures to resemble high-quality close-neighbor speech without much reverberation.

[0084] The augmentation module can be used by the data loader to generate training examples consisting of the augmented source as input, along with alternatively augmented versions of the same source as the target output. This allows for transformational examples where the target source can be represented in an alternatively augmented context when it is part of the training mixture. When used while training the modified RNN-CASSM, this allows the audio processing system to learn operations such as "filtering out," "dereverberation," and deeper restoration of highly obscured target sources.

[0085] Application examples 500 in the illustrated implementation include (i) filtering removal, which includes expanding the mixture with a filter; (ii) dereverberation, which includes expanding the mixture with a reverberation; (iii) background speaker recovery, which includes expanding the source with a filter and a reverberation; (iv) distortion repair, which includes expanding the mixture with a distortion; and (v) gap repair, which includes expanding the mixture with a gap.

[0086] With reference to Figures 6 and 7, an exemplary implementation of a machine learning training method 600 will now be described. In these examples, the machine learning training method will be described in relation to a method of training for improvements and / or useful functionality in output signal quality and source separation. With reference to block 3.1, a first machine learning training method includes a step of upscaling the trained network sample rate (e.g., from 24kHz to 48kHz). Due to limitations in time, computational resources, etc., the model is trained at 24kHz using a 24kHz dataset, but may be accompanied by limitations on output quality. An upscaling process undertaken on a model trained at 24kHz can provide functioning at 48kHz. One exemplary process includes a step of keeping the inner blocks of learned parameters of the masking network, while the encoder / decoder layers and those directly connected to them are discarded. In other words, only the inner separation layers are implanted into a newly initialized model with a 48kHz encoder / decoder. Next, the untrained 48kHz encoder / decoder and those directly connected layers are fine-tuned using the 48Khz dataset while the inherited network remains frozen. This is done until an acceptable validation / loss value (e.g., L1, L2, SISDR, SDSDR, or other loss calculation) is again identified during training / validation, which here denotes the adapted inherited layers. For example, the acceptable validation / loss value may be determined by observing a trend towards values ​​identified in previous training sessions compared to the model performance, by comparison to a predefined threshold loss value, or by other approaches. During training, the loss value ideally tends towards minimization, however, in practice, the loss value may also be useful in signaling significant problems during training, such as when the loss value starts to trend away from minimization. It has also been observed that while the loss value may not have improved during training, the model's performance as measured by the quality of source separation may still be improved by continuing training.

[0087] Finally, fine-tuning training continues across all layers, allowing the model to further evolve at 48Khz, with the end result being a well-performing 48Khz model. In some implementations, the system is trained to operate more quickly at high signal processing sample rates by training at a lower sample rate, inheriting appropriate layers into the higher sample rate model architecture, and performing a two-step training process in which, first, the untrained layers are trained while the parameters of the inherited layers are frozen, and, second, the entire model is then fine-tuned until its performance meets or exceeds that of the lower sample rate model. This process can be performed over many iterations, including, but not limited to, the following:

[0088] a) Train at 6kHz

[0089] b) Upscaling to 12kHz

[0090] c) Upscaling to 24kHz

[0091] d) Upscaling to 48kHz

[0092] Referring to block 3.2, multiple audio source mixtures are used to improve the performance of the speech isolation model (e.g., source=foreground, background, and distant speech, noise, and music mixtures). Speech initially trained with a single speech-to-noise / music mixture may not perform well, and the processed results may have problems consistently extracting from the original source media and suffer from substantial artifacting. Instead of training with a single speech-to-noise mixture, the audio processing system may provide substantially improved results by using multiple audio sources layered to simulate the variations in proximity of foreground and background speech, etc. This approach can also be applied to musical instruments, e.g., multiple overlaid guitar samples in a mixture instead of only one sample at a time. In various implementations, training samples with different layer scenarios may be selected to match and / or approximate the unidentified audio mixture (e.g., based on user input, identified source classes, and / or analysis of the unidentified audio mixture during iterative training)

[0093] The example training mixture 700 includes a speech isolation training mixture 702 that includes a mixture of sources and noise 704. The source mixture 706 may include a mixture of foreground speech 708, background speech 710, and far-field speech 712 in an expanded randomized combination. The noise mixture 714 includes a mixture of instruments 716, room tones 718, and hiss 720 in an expanded randomized combination.

[0094] With reference to Figures 8-10, an implementation of a machine learning process 800 will now be described in relation to a method of processing with a machine learning model that contributes to improvements and / or useful functionality in output signal quality and source separation (e.g., as previously discussed herein). With reference to block 4.1, the machine learning process may include a complement of the sum of the separated sources as an additional output. In other implementations, the model outputs speech and discarded music / noise. The complement output may subsequently be used in various processes to further process / separate the sources remaining in the complement output.

[0095] Referring to block 4.2, the machine learning post-processing model cleans up artifacts introduced by the machine learning process, such as clicks, harmonic distortion, ghosting, and broadband noise (artifacts may be determined, for example, as discussed above with reference to FIG. 4). The source separation model being trained may exhibit artifacts such as clicks, harmonic distortion, broadband noise, and "ghosting" where sounds are partially separated between the target complement outputs. In order for these outputs to be used in the context of a high quality soundtrack, a laborious cleanup would typically need to be attempted using conventional audio restoration software. Doing so may still result in undesirable qualities in the restored audio. This may be addressed by post-processing the processed audio with a model that has been trained on a dataset consisting of processing artifacts. The processing artifact dataset may be generated by the problematic model itself.

[0096] Once trained, the post-processing model 910 can be reused for all similar models. In the illustrated implementation, an input mixture 950 is processed with a generic model 952, which generates a source separation output with machine learning artifacts (step 954). A post-processing step 956 removes the artifacts and generates an enhanced output 960. The post-processing model 910 includes a step of generating a data set 914 of isolated machine learning artifacts (step 912). The machine artifacts may include clicks, ghosting, broadband noise, harmonic distortion, and other artifacts. The isolated machine learning artifacts 912 are used to train a model that removes the artifacts in step 916.

[0097] Referring to block 4.3, in some implementations, a user-guided self-iterative processing / training approach may be used. The user guides and contributes to the fine-tuning of the pre-trained model over the course of the processing / editing / training loop, which can then be used to bring about better source separation results than would have been possible from the general model. Model fine-tuning capabilities can be put in the hands of the user to solve source separations that the pre-trained model may not be able to solve. Processing with source separation models on unidentified mixtures is not always successful, usually due to a lack of sufficient training data. In one solution, the audio processing system uses a method whereby the user can guide and contribute to fine-tuning the pre-trained input, which can then be used to bring about better results. In an exemplary method, i) the user processes the input media, ii) is given the opportunity to assess the output or has the option to have this assessment performed by an algorithm that measures threshold parameters for some metrics (e.g., measured using a moving RMS window and / or other measurements as discussed above), iii) if the output is deemed acceptable, the processing ends here, otherwise iv) is given the opportunity to use temporal and / or spectral editing, and / or culling / expansion algorithms to manipulate the incomplete output. In essence, a partition of the output that will best serve the immediate step is selected. In step v), the media is now considered for inclusion in the training dataset, in step vi), the user is given the opportunity to also add their own auxiliary dataset, in step vii) the model is fine-tuned and trained according to the user's hyperparameter preferences, in step viii) the model's performance is verified to confirm the improved results from the previous iteration, and then in step ix) the process is repeated. In various implementations, hyperparameters associated with fine-tuning training may include parameters such as training segment length, epoch duration, training scheduler type and parameters, optimizer type and parameters, and / or other hyperparameters.The hyperparameters may be initially based on a set of predetermined values ​​and then modified by the user for fine-tuning training.

[0098] An exemplary user-guided self-iteration process 1000 is illustrated in FIG. 10. The process 1000 starts with a generic model 1002, which may be implemented as a pre-trained model as discussed above. An input mixture 1004, such as a single track audio signal with an unidentified source mixture or multiple single track audio signals, is processed through the generic model 1002 in step 1006 to generate separated audio signals from the mixture, including a machine learning separated source signal 1008 and a machine learning separated noise signal 1010. In step 1012, the results are evaluated to ensure that the separated audio sources have sufficient quality (e.g., by comparing the estimated MOS and a threshold and / or other quality measures as described above). If the results are determined to be good, the separated sources are output in step 1014.

[0099] If it is determined that the separated audio source signals require further improvement, one or more of the outputs are prepared for inclusion in a training dataset for fine-tuning in step 1016. In an automated fine-tuning system 1018, the machine learning separated sources 1008 and noise 1010 are used directly as a fine-tuning dataset 1034 (step 1022). The fine-tuning dataset is optionally culled in step 1036 based on a user-selected threshold of a source separation metric (e.g., comparing a moving RMS window to a threshold and / or other separation metrics as described above). Training is then performed in step 1038 to fine-tune the model. The fine-tuned model 1032 is then applied to the input mixture in step 1006.

[0100] In the user-guided fine-tuning system 1020, the user may select portions of the audio clip to include or omit from the fine-tuning dataset (step 1024-temporal editing). The user may also select portions of the clip to include / omit from the fine-tuning dataset and frequency / time window selections (step 1026-spectral editing). In some implementations, the user may provide additional audio clips to extend the fine-tuning dataset (step 1028-add to dataset). In some implementations, the user provides equalization, reverb, distortion, and / or other enhancement settings to fine-tune the dataset (step 1030-enhance). After the fine-tuning dataset 1034 is updated, the training process continues with steps 1036-1038 to generate a fine-tuned model 1032.

[0101] Referring to block 4.4, an animated visual representation of the model fine-tuning progress may be implemented. While the user fine-tunes the model to solve source separation for a particular media clip, the model's progressive output is displayed to help indicate the model's performance to help guide the user's decision-making. In some implementations, for example, an interface may be displayed in a window with associated tool icons that displays a periodically updated spectrogram animated representation of the estimated outputs as computed by the fine-tuned model, which is periodically tested while it is being fine-tuned. This interface may allow the user to visually assess how well the model is currently performing in various regions of the input mixture. The interface may also facilitate user interaction, such as allowing the user to experiment with time / frequency selection of these estimated outputs based on interaction with the spectrogram window.

[0102] Referring to block 4.5, user-guided extensions for fine-tuning training may be implemented. To improve results when enhancing / separating targeted recordings with specific characteristics such as reverberation, filtering, nonlinear distortion, etc., the audio processing system may present the user with tools to control the underlying algorithms for guiding and contributing to the selection of the extensions, such as reverberation, filtering, nonlinear distortion, and / or noise, during step 1020 of the loop described in block 4.3. The user can guide and contribute to the selection of the extensions and / or extension parameters used during the fine-tuning of the pre-trained model over the course of the processing / editing / training loop. The extensions include user-controllable randomization settings for each parameter (e.g., values ​​affecting various aspects of the extension, such as the intensity, density, modulation, and / or decay of the reverberation extension) to help generalize or narrow the targeted behavior after fine-tuning. This allows for an expansion of the control that users have over fine-tuning training, which allows them to specifically match, for example, the filtered sounds / reverberation in the targeted recordings. When referring to a random process or randomization, it may be sufficient to have a pseudo-random process or an arbitrary selection process. In some implementations, the extension parameters are automatically matched to the input source mixture using one or more algorithms. In some implementations, the automatic match is used to achieve a rough match as a starting point. For example, a spectral analysis of the target input mixture may be combined with an analysis of the random dataset samples to yield a set of frequency band deviation scores that can be used to minimize deviations between the random extension dataset samples and the target input mixture by adjusting various parameters of the dataset extension filters based on the values ​​of the frequency band deviation scores.

[0103] 11A and 11B, an exemplary implementation of a machine learning application 1100 will now be described according to one or more implementations. The audio processing systems and methods disclosed herein may be used in conjunction with other audio processing applications that contribute improvements and / or useful functionality to sound post-production editing workflows.

[0104] Referring to block 5.1, a plug-in such as the Avid Audio Extension (AAX) plug-in hosted in a digital audio workstation (DAW) (e.g., a digital audio workstation sold under the trade name PRO TOOLS may be used in one or more implementations) is provided that allows a user to send audio clips from a standalone application where they may subsequently be processed using a machine learning model. The plug-in can return an arbitrary number of stem splits back to the DAW environment.

[0105] Referring to block 5.2, implementations of the present disclosure also load / receive media used in an application (e.g., an application sold under the trade name JAM LAB with JAM CONNECT or a similar application) that is then processed by a user-selected machine learning recipe. In some implementations, multiple client machines 1 102A-C and processing nodes 1 106A-D having access to client software (e.g., JAM LAB with JAM CONNECT) are configured to access a task manager / database 1104 for access to both the client software and machine learning (ML) application 1100 as disclosed herein.

[0106] With reference to FIG. 12, an exemplary process flow will now be described. A system running a digital audio workstation 1202 (e.g., PRO TOOLS with JAM CONNECT) is configured to send audio clips to a client application 1208 (e.g., including JAM LAB). Through the client application, a single model may be selected from a list of categories 1210 to process / separate audio sources from the audio clip. A type of stem is also selected in step 1212 to form a multi-model recipe. The audio clip and recipe are sent (in step 1214) to a task manager / database 1216, which manages and distributes the recipe processing across the user's available processing nodes in step 1218. The client application receives and returns the processed audio clip labeled with the stem and / or model name in step 1206. In some implementations, the client application 1208 may also facilitate selecting stems to form a multi-model recipe 1212 as disclosed herein.

[0107] An implementation of a machine learning multi-model recipe processing system will now be described with reference to Figures 13A-E. In some implementations, a user may desire to separate targeted media into a set of source classes / stems in one step using a selection of source separation models in a particular order and hierarchical combination. A sequential / divergent source separation recipe schema is implemented to process targeted media using one or more source separation models in a sequential / divergent structured order to separate the targeted media into a set of user-selected source classes or stems.

[0108] In the implementation 1300 of FIG. 13A, each step of the recipe represents a source separation model or combination of models targeting a particular source class. The recipe includes models defined and appropriately trained according to the recipe schema for performing the steps that result in the user's desired stem output. As illustrated in the example implementation, the user may choose to first separate hiss, followed by voice (which is then separated into vocals and other speech), drums (which are then separated into kick, snare, and other percussion), organ, piano, bass, and other processing. Thus, by processing the targeted media through a pipeline defined by the recipe, the output at the various steps may be collected, ultimately separating the targeted media into a set of user-selected source classes or stems.

[0109] Referring to Figure 13B, an example voice / drums etc. sequential processing pipeline 1320 is illustrated. In this model, the input mixture is first processed to extract the voice, and the complement includes the drums and other sounds in the mixture. The drum model extracts the drums, and the complement includes the other sounds. In this implementation, the output includes the voice, drums and other stems. The order in which models are applied when separating source classes using a sequential / divergent separation system may be optimized for higher quality by using an algorithm that assesses the optimal processing order. An exemplary implementation of the optimized processing method 1340 is illustrated in FIG. 13C. For example, an input mixture model with A and B components may be configured to isolate class A and then the remaining B. This may yield different results if the processing order is reversed (e.g., isolate B and then A). The optimized processing method 1340 may operate on an input mixture of A+B by separating A and B in both orders, comparing the results, and selecting the order with the best results (e.g., the results with fewer errors in the separated stems). In various implementations, the optimized processing method 1340 may operate manually, automatically, and / or in a hybrid approach. For example, an estimated optimal order may be pre-established by using a set of ground truth test samples and then creating a test mixture that can be separated using various permutations of stem order, and then an estimated best performing model processing order may be established by using an error function that compares the output of these tests against the ground truth test samples.

[0110] An exemplary pipeline 1360 for improving the output fidelity of a sequential / divergent source separation system (e.g., as described previously herein) will now be described with reference to FIG. 13D. For example, the output fidelity may be measured using the MOS algorithm, which is capable of measuring the deviation of the output of a stem from a particular labeled dataset, such as a collection of speech samples from an individual. In some implementations, such an algorithm may be implemented as a neural network that is pre-trained to either classify sources or measure the deviation of sources from a given dataset.

[0111] A processing recipe (see 5.2.1) may also include a post-processing step after one or more outputs in the pipeline. The post-processing step may include any type of digital signal processing filter / algorithm that cleans up signal artifacts / noise that may have been introduced by prior steps. The pipeline 1360 uses a model specifically trained to clean up artifacts (e.g., as described in step 4.2 herein), resulting in greatly improved overall results, especially due to the sequential nature of the recipe processing pipeline.

[0112] 13E, an example implementation 1380 combines models to separate specific source classes. Sequential / divergent separation processing recipes (see 5.2.1) may include steps where one or more models are used in combination to extract source classes from the mixture that could not otherwise be fully extracted by only one model trained to target that source class. In the illustrated example, drums are separated twice, subsequently summed, and presented as a single "drum" output stem, accompanied by a complementary "other" stem.

[0113] An exemplary audio processing system 1400 for implementing the systems and methods disclosed herein will now be described with reference to Figure 14. The audio processing system 1400 includes a logic device 1402, a memory 1404, a communication component 1422, a display 1418, a user interface 1420, and a data storage device 1430.

[0114] The logic device 1402 may include, for example, a microprocessor, a single core processor, a multi-core processor, a microcontroller, a programmable logic device configured to perform processing operations, a DSP device, one or more memories for storing executable instructions (e.g., software, firmware, or other instructions), a graphics processing unit, and / or any other suitable combination of processing devices and / or memories configured to execute instructions to perform any of the various operations described herein. The logic device 1402 is adapted to interface and communicate with various components of the audio processing system 1400, including the memory 1404, the communication component 1422, the display 1418, the user interface 1420, and the data storage 1430.

[0115] The communications component 1422 may include wired and wireless communications interfaces to facilitate communication with a network or remote systems. The wired communications interfaces may be implemented as one or more physical network or device connection interfaces, such as a cable or other wired communications interfaces. The wireless communications interfaces may be implemented as one or more Wi-Fi, Bluetooth, cellular, infrared, radio, and / or other types of network interfaces for wireless communication. The communications component 1422 may include an antenna for wireless communication during operation.

[0116] The display 1418 may include an image display device (e.g., a liquid crystal display (LCD)) or various other types of commonly known video displays or monitors. The user interface 1420, in various implementations, may include user input and / or interface devices, such as a keyboard, a control panel unit, a graphical user interface, or other user input / output. The display 1418 may operate as both a user input device and a display device, such as, for example, a touch screen device adapted to receive input signals from a user touching different portions of a display screen.

[0117] Memory 1404 stores program instructions for execution by logic device 1402, including program logic for implementing the systems and methods disclosed herein, including, without limitation, audio source separation tools 1406, core model operations 1408, machine learning training 1410, trained audio separation models 1412, audio processing applications 1414, and self-iterative processing / training logic 1416. Data used by audio processing system 1400 may be stored in memory 1404 and / or in data storage 1430 and may include machine learning speech dataset 1432, machine learning music / noise dataset 1434, audio stems 1436, audio mixtures 1438, and / or other data.

[0118] In some implementations, one or more processes may be implemented through a remote processing system, such as a cloud platform, which may be implemented as audio processing system 1400 as described herein.

[0119] FIG. 15 illustrates an example neural network that may be used in one or more of the implementations of FIGS. 1-14, including various RNNs and models as described herein. The neural network 1500 is implemented as a recurrent neural network, a deep neural network, a convolutional neural network, or other suitable neural network that receives a labeled training data set 1510 to generate an audio output 1512 (e.g., one or more audio stems) for each input audio sample. In various implementations, the labeled training data set 1510 may include various audio samples and training mixtures as described herein, such as a training data set (FIG. 3), an auto-iterative training data set or a culled data set (FIG. 4), training mixtures and data sets described according to the training methods described herein (FIGS. 5-13E), or other training data sets, as appropriate.

[0120] The training process for generating a trained neural network model includes a forward pass through the neural network 1500 to generate an audio stem or other desired audio output 1512. Each data sample is labeled with the desired output of the neural network 1500, which is compared to the audio output 1512. In some implementations, a cost function may be applied to quantify the error in the audio output 1512, and a backward pass through the neural network 1500 may then be used to adjust the neural network coefficients to minimize the output error.

[0121] The trained neural network 1500 may then be tested for accuracy using a subset of the labeled training data 1510 reserved for validation. The trained neural network 1500 may then be implemented as a model in a runtime environment to perform audio source separation as described herein.

[0122] In various implementations, the neural network 1500 processes input data (e.g., audio samples) using an input layer 1520. In some examples, the input data may correspond to audio samples and / or audio inputs as previously described herein.

[0123] The input layer 1520 includes a number of neurons used to condition the input audio data for input to the neural network 1500, which may include feature extraction, scaling, sampling rate conversion, and / or the like. Each neuron in the input layer 1520 generates an output that feeds into the input of one or more hidden layers 1530. The hidden layers 1530 include a number of neurons that process the output from the input layer 1520. In some embodiments, each neuron in the hidden layer 1530 generates an output that is collectively then propagated through additional hidden layers that include a number of neurons that process the output from a prior hidden layer. The output of the hidden layer 1530 is fed to the output layer 1540. The output layer 1540 includes one or more neurons used to condition the output from the output layer 1540 to generate a desired output. It should be understood that the architecture of neural network 1500 is merely representative and that other architectures are possible, including neural networks with only one hidden layer, neural networks with no input and / or output layers, neural networks with recurrent layers, and / or the like.

[0124] In some embodiments, the input layer 1520, the hidden layer 1530, and / or the output layer 1540 each include one or more neurons. In some embodiments, the input layer 1520, the hidden layer 1530, and / or the output layer 1540 may each include the same or different numbers of neurons. In some embodiments, each neuron takes a combination (e.g., a weighted sum using a trainable weighting matrix W) of its inputs x, adds an optional trainable bias b, and applies an activation function f to generate an output α as shown in the equation α=f(Wx+b). In some embodiments, the activation function f may be a linear activation function, an activation function with upper and / or lower bounds, a log-sigmoid function, a hyperbolic tangent function, a rectified linear unit function, and / or the like. In some embodiments, each neuron may have the same or different activation functions.

[0125] In some embodiments, the neural network 1500 may be trained using supervised learning, where combinations of training data include combinations of input data and ground truth (e.g., expected) output data. Differences between the generated audio output 1512 and the ground truth output data (e.g., labels) are fed back into the neural network 1500 to correct for various trainable weights and biases. In some embodiments, the differences may be fed back using backpropagation techniques using a stochastic gradient descent algorithm and / or the like. In some embodiments, a large set of training data combinations may be presented to the neural network 1500 multiple times until an overall cost function (e.g., mean squared error based on the differences of each training combination) converges to an acceptable level.

[0126] An example implementation is described below.

[0127] 1. An audio processing system comprising a deep neural network (DNN) trained to separate one or more audio source signals from a single track audio mixture.

[0128] 2. The audio processing system of Example 1, wherein the DNN is configured to receive a signal input and generate a signal output without time-domain encoding and / or time-domain decoding.

[0129] 3. The audio processing system of Examples 1-2, wherein the DNN is configured to apply a window function.

[0130] 4. The audio processing system of Examples 1-3, wherein the DNN performs an overlap-add process to smooth out banding artifacts.

[0131] 5. The audio processing system of any one of Examples 1-4, wherein audio source separation is performed without applying a mask.

[0132] 6. The audio processing system described in Examples 1-5, where the DNN model is trained using a 48 kHz sample rate.

[0133] 7. The audio processing system of any one of Examples 1-6, wherein the signal processing pipeline operates at 48 kHz.

[0134] 8. An audio processing system as described in any one of Examples 1-7, further comprising a separation strength parameter that controls the strength of the separation process applied to the input audio signal.

[0135] 9. The audio processing system of any one of Examples 1-8, further comprising a speech training dataset comprising a plurality of labeled speech samples.

[0136] 10. The audio processing system of any one of Examples 1-9, further comprising a non-speech training data set comprising a plurality of labeled music and / or noise data samples.

[0137] 11. An audio processing system as described in Examples 1-10, further comprising a dataset generation module configured to generate labeled audio samples for use in training a DNN model.

[0138] 12. The audio processing system of any of Examples 1-11, wherein the data set generation module is a self-iterating data set generator.

[0139] 13. The audio processing system of Examples 1-12, wherein the dataset generation module is configured to generate labeled audio samples from the input audio mixture and / or the audio source stems output from the DNN.

[0140] 14. The audio processing system of any one of Examples 1-13, further comprising a data loader configured to apply pre / post blend enhancements.

[0141] 15. The audio processing system of Examples 1-14, wherein the DNN is trained at higher than audible frequencies to recognize distinct stems of audio in the lower audible frequency range.

[0142] 16. The audio processing system of any of Examples 1-15, wherein the data loader is configured to apply enhancements such as reverb, filters, probability parameters, etc.

[0143] 17. The audio processing system of Examples 1-16, further configured to match the sound quality of the target unidentified mixture during training based on the relative signal level.

[0144] 18. An exemplary method, comprising:

[0145] processing the audio input data using a trained inference model trained for source separation to generate source-separated stems;

[0146] generating a speech dataset from the source separated stem;

[0147] generating a noisy data set from a source separation system;

[0148] training the inference model using the speech dataset and the noise dataset to generate an updated inference model; A method comprising:

[0149] 19. The method of example 18, further comprising processing audio input data using the updated inference model.

[0150] 20. The method of Examples 18-19, further comprising the step of iteratively updating the updated inference model.

[0151] 21. The method of Examples 19-20, wherein the training dataset is curated to include samples that approximate the audio source.

[0152] 22. The method of any one of Examples 19-21, further comprising a hierarchical mix bus schema.

[0153] 23. The method of Examples 19-22, in which the inference model is trained using a multiple audio source mixture.

[0154] 24. The method of any of Examples 19-23, wherein the inference model is trained using foreground speech, background speech, and / or distant speech.

[0155] 25. The method of any one of Examples 19-24, wherein the inference model is trained at a first sample rate and upscaled to a higher sample rate.

[0156] 26. The method of any of Examples 19-25, further comprising the step of post-processing the separated audio source stems to remove artifacts introduced by the source separation process.

[0157] 27. The method of any one of Examples 19-26, wherein the source separation system includes a separated source signal and a remaining complementary signal.

[0158] 28. The method of any one of Examples 19-27, wherein the artifacts introduced during the source separation process include clicks, harmonic distortion, ghosting, and / or broadband noise.

[0159] 29. The method of any one of Examples 19-28, wherein the fine-tuning process includes a user-guided self-iterative process.

[0160] 30. The method of any one of Examples 19-29, further comprising facilitating user-guided enhancements for fine-tuning of training.

[0161] 31. A system comprising: an audio input configured to receive an audio input stream comprising a mixture of audio signals generated from a plurality of audio sources; a trained audio source separation model configured to receive the audio input stream and generate a plurality of generated audio stems, the generated plurality of audio stems corresponding to one or more of the plurality of audio sources; and a self-iterative training system configured to update the trained audio source separation model to an updated audio source separation model based, at least in part, on the generated plurality of audio stems, the updated audio source separation model configured to reprocess the audio input stream and generate a plurality of improved audio stems.

[0162] 32. The system of Example 31, wherein the audio input stream comprises one or more single-track audio mixtures, and the trained audio source separation model comprises a neural network that is trained to separate one or more audio source signals from the one or more single-track audio mixtures.

[0163] 33. The system of Examples 31-32, wherein the neural network is configured to perform audio source separation without applying a mask.

[0164] 34. The system of Examples 31-33, further comprising a training dataset comprising labeled source audio data and labeled noisy audio data, wherein the trained audio source separation model is trained using the training dataset to generate a general source separation model.

[0165] 35. The system of Examples 31-34, wherein at least a subset of the generated plurality of audio stems is culled based on a threshold metric and added to a training dataset to form a culled dynamically evolving dataset, and the culled dynamically evolving dataset is used to train an updated audio source separation model.

[0166] 36. The system described in Examples 31-35, wherein the self-iterative training system is further configured to calculate a first quality metric associated with the generated plurality of audio stems, the first quality metric providing a first performance measure of the trained audio source separation model, and the self-iterative training system is further configured to calculate a second quality metric associated with the improved audio stems, the second quality metric providing a second performance measure of the updated audio source separation model, the second quality metric exceeding the first quality metric.

[0167] 37. The system of Examples 31-36, wherein the trained audio source separation model is trained using a training dataset comprising a plurality of datasets, each of the plurality of datasets comprising labeled audio samples configured to train the system to address a source separation problem.

[0168] 38. The system of Examples 31-37, wherein the multiple datasets comprise a speech training dataset comprising a plurality of labeled speech samples and / or a non-speech training dataset comprising a plurality of labeled music and / or noise data samples.

[0169] 39. The system of any one of Examples 31-38, wherein the self-repetitive training system further comprises a self-repetitive data set generation module configured to generate labeled audio samples from the generated plurality of audio stems.

[0170] 40. The system of Examples 31-39, wherein the multiple enhanced audio stems are generated using a hierarchical divergent sequence that includes a step of separating the source signal and the remaining complementary signal.

[0171] 41. A method comprising: receiving an audio input stream comprising a mixture of audio signals generated from a plurality of audio sources; using a trained audio source separation model configured to receive the audio input stream to generate a plurality of generated audio stems corresponding to one or more of the plurality of audio sources; updating the trained audio source separation model to an updated audio source separation model based, at least in part, on the generated plurality of audio stems using a self-iterative training process; and reprocessing the audio input stream using the updated audio source separation model to generate a plurality of improved audio stems.

[0172] 42. The method of Example 41, wherein the audio input stream comprises one or more single-track audio mixtures, and the trained audio source separation model comprises a neural network that is trained to separate one or more audio source signals from the one or more single-track audio mixtures.

[0173] 43. The method of any one of Examples 41-42, wherein the neural network is configured to perform audio source separation without applying a mask.

[0174] 44. The method of any one of Examples 41-43, further comprising the steps of providing a training dataset comprising labeled source audio data and labeled noise audio data, and using the training dataset to train an audio source separation model trained to generate a general source separation model.

[0175] 45. The method of any one of Examples 41-44, further comprising the steps of adding at least a subset of the generated plurality of audio stems to a training dataset to generate a dynamically evolving dataset, culling the dynamically evolving dataset based on a threshold metric, and training an updated audio source separation model using the culled dynamically evolving dataset.

[0176] 46. ​​The method of any one of Examples 41-45, wherein the self-iterative training process further includes the steps of calculating a first quality metric associated with the generated plurality of audio stems, the first quality metric providing a first performance measure of the trained audio source separation model; calculating a second quality metric associated with the improved audio stems, the second quality metric providing a performance measure of the updated audio source separation model; and comparing the second quality metric with the first quality metric and confirming that the second quality metric exceeds the first quality metric.

[0177] 47. The method described in Examples 41-46, wherein the trained audio source separation model is trained using a training dataset comprising a plurality of datasets, each of the plurality of datasets comprising labeled audio samples configured to train the audio source separation model to address a different source separation problem.

[0178] 48. The method of any one of Examples 41-47, wherein the multiple datasets comprise a speech training dataset comprising a multiple labeled speech samples and / or a non-speech training dataset comprising a multiple labeled music and / or noise data samples.

[0179] 49. The method described in Examples 41-48, wherein the self-repetitive training process further includes a step of generating labeled audio samples from a plurality of audio stems generated for the self-repetitive data set.

[0180] 50. The method of any one of Examples 41-49, further comprising generating a plurality of enhanced audio stems using a hierarchical divergent sequence, including separating the source signal and the remaining complement signal.

[0181] Where applicable, the various implementations provided by the present disclosure can be implemented using hardware, software, or a combination of hardware and software. Also, where applicable, the various hardware and / or software components described herein can be combined into composite components comprising software, hardware, and / or both without departing from the spirit of the present disclosure. Where applicable, the various hardware and / or software components described herein can be separated into subcomponents comprising software, hardware, or both without departing from the spirit of the present disclosure.

[0182] Software according to the present disclosure, such as non-transient instructions, program code, and / or data, can be stored on one or more non-transient machine-readable media. It is also envisioned that the software identified herein can be implemented using one or more general-purpose or special-purpose computers and / or computer systems, networked and / or otherwise. Where applicable, the ordering of the various steps described herein can be changed, combined into composite steps, and / or separated into sub-steps to provide the features described herein. The implementations described above illustrate the invention but do not limit it. It should also be understood that numerous modifications and variations are possible in accordance with the principles of the invention. The scope of the invention is therefore defined solely by the following claims.

Claims

1. 1. A method, comprising: receiving a single track audio input sample comprising an unknown mixture of audio signals generated from multiple audio sources; Separating one or more of the audio sources from the single track audio input samples using a sequential audio source separation model; Including, The separating step comprises: (a) defining a processing recipe including a plurality of source separation processes, the plurality of source separation processes including at least a first source separation process and a second source separation process, the first source separation process configured to process an audio input mixture and output a first set of one or more separated source signals and a first complementary residual signal mixture, and the second source separation process configured to process the first complementary residual signal mixture and output a second set of one or more separated source signals and a second complementary residual signal mixture; (b) generating a plurality of audio stems separated from the unknown mixture of audio signals by processing the single track audio input samples according to the processing recipe; A method comprising:

2. Defining the process recipe comprises: generating a first set of audio stems by processing the single track audio input samples using a first processing order of the multiple source separation processes to separate audio stems; generating a second set of audio stems by processing the single track audio input samples using a second processing order of the multiple source separation processes that is different from the first processing order for separating audio stems; evaluating the first set of audio stems and the second set of audio stems to determine which of the first processing order and the second processing order should be included in the processing recipe; The method of claim 1 , comprising:

3. 2. The method of claim 1 , wherein the processing recipe further comprises post-processing the multiple audio stems to remove and / or mitigate artifacts introduced by the sequential audio source separation model, wherein the post-processing comprises processing the multiple audio stems through one or more neural networks trained to remove and / or mitigate artifacts comprising clicks, harmonic distortion, ghosting, and / or broadband noise.

4. The method of claim 1, wherein at least one of the first set of one or more separated source signals and the second set of one or more separated source signals comprises a mixture of a speech source and a corresponding one of the first remaining complementary signal mixture or the second remaining complementary signal mixture comprising music and noise.

5. The process recipe further comprises a sequential branching process sequence including one or more subsequent separation processes; 2. The method of claim 1, wherein each subsequent separation process processes an output signal comprising an additional set of one or more separated source signals or an additional remaining complementary signal mixture from a previous source separation process, and is configured to separate a particular source class from the processed output signal.

6. A method, comprising: receiving a single track audio input sample comprising an unknown mixture of audio signals generated from multiple audio sources; Separating one or more of the audio sources from the single track audio input samples using a sequential audio source separation model; Including, The separating step comprises: (a) defining a plurality of process recipes, each process recipe including a plurality of source separation processes having a corresponding order of execution; (i) each of the plurality of source separation processes is configured to receive an audio input mixture and generate a plurality of audio output stems comprising one or more separated source signals for a corresponding source class and a remaining complementary signal mixture; (ii) the audio input mixture for a first one of the source separation processes in the order of execution comprises the single track audio input samples, and the audio input mixture for each subsequent source separation process comprises either the one or more separated source signals from a preceding source separation process or the remaining complementary signal mixture; And, (b) generating a corresponding set of audio output stems separated from the unknown mixture of audio signals by processing the single track audio input sample using each of the plurality of processing recipes; (c) selecting one of the plurality of processing recipes for the sequential audio source separation model by evaluating the set of audio output stems from the plurality of processing recipes to determine the processing recipe that generates the highest quality set of audio output stems compared to properties of the corresponding source class; A method comprising:

7. the processing recipe includes post-processing the plurality of audio stems; The post-processing step comprises: combining two or more of the plurality of audio stems having a common source class that were generated by different source separation processes within the processing recipe; and / or processing one or more of the plurality of audio stems to mitigate artifacts and / or noise introduced by the sequential audio source separation model by applying a neural network model trained to clean up artifacts on the audio stems separated according to the processing recipe. The method of claim 1 , comprising:

8. wherein separating one or more of the audio sources from the single track audio input samples using a sequential audio source separation model further comprises a user-guided and / or self-iterative process configured to progressively separate sources from the unknown mixture of audio signals; said user-guided and / or self-iterative process comprising: (a) selecting audio enhancements and / or audio enhancement parameters to fine-tune the sequential audio source separation model to enable matching of source signals in the unknown mixture of audio signals; (b) identifying a generic separation model for use in the processing recipe by identifying source classes within the unknown mixture of audio signals; The method of claim 1 , comprising:

9. the processing recipe comprises a set of source classes and output stems and corresponding source separation models arranged in a hierarchical branching sequence; The method of claim 8 , wherein at least one class of sources is separated into a plurality of source class separation stems, and the processing recipe includes generating a mixture of the plurality of source class separation stems as a source class output.

10. The method further includes training the sequential audio source separation model using the single track audio input sample by providing the single track audio input sample to a generic source separation model to generate a plurality of initial audio stems corresponding to one or more of the plurality of audio sources, and retraining the generic source separation model at least in part using one or more of the plurality of initial audio stems; 2. The method of claim 1 , wherein the general source separation model comprises a plurality of neural network models, each configured to process a single-channel audio input sample and output one or more source-separated audio stems comprising a complementary signal mixture of a source class and a residual.

11. training the sequential audio source separation models further comprises training the plurality of neural network models in an order defined by the processing recipe; The method of claim 10 , wherein the processing recipe comprises a hierarchical branching sequence, each branch comprising one or more of the plurality of neural network models.

12. 12. The method of claim 11 , wherein training the sequential audio source separation model further comprises re-iteratively training the sequential audio source separation model based at least in part on the plurality of audio stems generated from prior iterations of the sequential audio source separation model.

13. Training the sequential audio source separation model includes: evaluating the one or more of the plurality of audio stems by determining a metric associated with the source-separated audio stem and comparing the metric to one or more threshold parameters; adding one or more source-separated audio stems to a training dataset to train the sequential audio source separation model based on said evaluating; and The method of claim 12 further comprising:

14. The method of claim 1, wherein training further comprises processing the source separation output with artifacts by generating a training dataset of artifacts generated by the processing recipe to train one or more neural networks, and generating an improved output in which one or more artifacts are reduced and / or removed in accordance with the processing recipe.

15. 1. A system comprising: a memory component that stores machine-readable instructions; Logical Devices and Equipped with The logic device executes the machine-readable instructions to: (a) separating one or more audio sources from a single-track audio input sample comprising an unknown mixture of audio signals generated from multiple audio sources using a sequential audio source separation model; and The separating step comprises: (1) defining a processing recipe including a plurality of source separation processes, the plurality of source separation processes including at least a first source separation process and a second source separation process, the first source separation process configured to process an audio input mixture and output one or more separated source signals and a first complementary residual signal mixture, and the second source separation process configured to process the first complementary residual signal and output one or more separated source signals and a second complementary residual signal mixture; (2) generating a plurality of audio stems separated from the unknown mixture of audio signals by processing the single track audio input samples according to the processing recipe; The system is implemented by

16. The logical device is generating a first set of audio stems by processing the single track audio input samples using a first processing order of the multiple source separation processes to separate audio stems; generating a second set of audio stems by processing the single track audio input samples using a second processing order of the multiple source separation processes that is different from the first processing order for separating audio stems; evaluating the first set of audio stems and the second set of audio stems to determine which of the first processing order and the second processing order should be included in the processing recipe; The system of claim 15 , further configured to define the process recipe by:

17. The logical device is and further configured to define a processing recipe by post-processing the plurality of audio stems to remove and / or mitigate artifacts introduced by the sequential audio source separation model; 16. The system of claim 15, wherein the post-processing comprises processing the multiple audio stems through one or more neural networks trained to remove and / or mitigate artifacts comprising clicks, harmonic distortion, ghosting, and / or broadband noise.

18. The system described in claim 15, wherein at least one of the first set of one or more separated source signals and the second set of one or more separated source signals comprises a mixture of a speech source and a corresponding one of the first remaining complementary signal mixture or the second remaining complementary signal mixture comprising music and noise.

19. The method of claim 18, wherein the process recipe further comprises a sequential branching process sequence including one or more subsequent separation processes.

16. The system of claim 15, wherein each subsequent separation process processes an output signal comprising an additional set of one or more separated source signals or an additional remaining complementary signal mixture from a previous source separation process, and is configured to separate a particular source class from the processed output signal.

20. The logical device: defining a plurality of process recipes, each process recipe having a different corresponding order of execution for the plurality of source separation processes; processing the single track audio input sample with each of the plurality of processing recipes to generate a corresponding set of audio output stems separated from the unknown mixture of audio signals; selecting one of the plurality of processing recipes for the sequential audio source separation model by evaluating the set of audio output stems from the plurality of processing recipes to determine the processing recipe that generates the highest quality set of audio output stems compared to properties of a corresponding source class; The system of claim 15 , further configured to: