Audio source separation processing pipeline system and method
The system uses machine learning models to separate and enhance audio components from low-quality single-track recordings, effectively addressing the challenge of generating high-fidelity stems by refining speech stems and reducing noise artifacts.
Patent Information
- Application Number
- JP2025180213
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-23
- Filing Date
- 2025-10-27
- Publication Date
- 2026-02-03
AI Technical Summary
Existing audio source separation techniques are not optimized to generate high-quality audio stems from low-quality, single-track noisy sound mixtures, particularly from older sound recordings, which are common in music and film industries.
A system and method using machine learning models to separate and enhance audio components from single-track recordings, involving sequential audio source separation models, neural networks for artifact removal, and a self-iterative training process to refine speech stems and mitigate noise.
The system effectively separates and enhances audio components into high-fidelity stems, addressing the challenges of low-quality recordings by improving audio quality and reducing artifacts such as clicks, harmonic distortion, and broadband noise.
Smart Images

Figure 2026016593000001_ABST
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This disclosure claims priority to U.S. Patent Application No. 17 / 848,341, filed June 23, 2022, and claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 272,650, filed October 27, 2021, both of which are incorporated by reference herein in their entireties.
[0002] The present disclosure relates generally to systems and methods for audio source separation, and more particularly to systems and methods for separating and enhancing audio source signals from an audio mixture, such as a single-track audio mixture. [Background technology]
[0003] Audio mixing is the process of combining multiple audio recordings to produce an optimized mixture for playback in one or more desired sound formats, such as mono, stereo, or surround sound. In applications requiring high-quality sound production, such as sound production for music and movies, audio mixtures are typically produced by mixing separate, high-quality recordings. These separate recordings are often generated in a controlled environment, such as a recording studio, with optimized sound effects and high-quality recording equipment.
[0004] Often, some of the source audio is of poor quality and / or may contain a mixture of desired audio sources and unwanted noise. In modern audio post-production, it is common to re-record audio when the original recording lacks the desired quality. For example, in music recording, vocal or instrument tracks may be recorded and mixed with the previous recording. In sound post-production for films, it is common to bring actors into a studio to re-record their lines and add other audio (e.g., sound effects, music) to the mix.
[0005] However, in some applications, it is desirable to faithfully convert the original audio source into a high-quality audio mix. For example, movies, music, television broadcasts, and other audio recordings can date back more than 100 years. The source audio was recorded on older, lower-quality equipment and may contain a low-quality mixture of desired audio and noise. For many recordings, a single-track / mono audio mix is the only audio source available to create an optimized mix for playback on modern sound systems.
[0006] One approach to processing an audio mixture is to separate the audio mixture into a set of separate audio source components and generate a separate audio stem for each component of the audio mixture. For example, a music recording may be separated into a vocal component, a guitar component, a bass component, and a drum component. Each of the separate components may then be enhanced and mixed to optimize playback.
[0007] However, existing audio source separation techniques are not optimized to generate the high-quality audio stems needed to produce high-fidelity output for the music and film industries. Audio source separation is particularly difficult when the audio sources are low-quality, single-track noisy sound mixtures from older sound recordings.
[0008] In view of the foregoing, there is a continuing need for improved audio source separation systems and methods, particularly for the generation of high fidelity audio from lower quality audio sources.
[0009] It is an object of at least the preferred embodiments to address at least some of the aforementioned disadvantages. An additional or alternative object is to at least provide the public with a useful alternative to conventional techniques. Summary of the Invention [Means for solving the problem]
[0010] Improved audio source separation systems and methods are disclosed herein. In various implementations, a single-track audio recording is provided to an audio source separation system configured to separate and separate various audio components, such as speech and individual instruments, into high-fidelity stems (e.g., discrete or grouped collections of audio sources mixed together).
[0011] In some implementations, an audio source separation system includes a first machine learning model that is trained to separate a single-track audio recording into stems, including speech, complements, and artifacts such as "clicks." Additional machine learning models may then be used to refine the speech stems by removing processing artifacts from the speech and / or fine-tuning the first machine learning model.
[0012] The term "comprising" as used herein means "consisting at least in part of." When interpreting each phrase herein containing the term "comprising," features other than the one or more prefaced by the term may also be present. Related terms such as "comprise" and "comprises" are to be interpreted in the same manner.
[0013] In various implementations, a method includes receiving a single-track audio input sample comprising an unknown mixture of audio signals generated from multiple audio sources; separating one or more of the audio sources from the single-track audio input sample using a sequential audio source separation model, the method including: defining a processing recipe including a multiple source separation process configured to receive the audio input mixture and output one or more separated source signals and a remaining complementary signal mixture; and processing the single-track audio input sample according to the processing recipe to generate multiple audio stems separated from the unknown mixture of audio signals.
[0014] The method may further define a processing recipe by processing the single track audio input samples using a first processing order of the multiple source separation processes for separating the audio stems to generate a first set of audio stems, processing the single track audio input samples using a second processing order of the multiple source separation processes that is different from the first processing order for separating the audio stems to generate a second set of audio stems, and evaluating the first set of audio stems and the second set of audio stems to determine which of the first processing order and the second processing order should be included in the processing recipe.
[0015] The processing recipe of the method may further include post-processing the multiple audio stems to remove and / or mitigate artifacts introduced by the sequential audio source separation model, where the post-processing includes processing the multiple audio stems through one or more neural networks trained to remove and / or mitigate artifacts comprising clicks, harmonic distortion, ghosting, and / or broadband noise.
[0016] The processing recipe of the method may further include executing multiple source separation processes in a sequential branching processing order, where a first source separation process is configured to receive a single track audio input sample and output one or more source separation signals and a remaining complementary signal mixture, and where each subsequent separation process is configured to receive an output signal from a previous source separation process, and where each node of the processing recipe is configured to separate a particular source class from the input audio mixture.
[0017] The processing recipe of the method may further include post-processing the plurality of audio stems by combining two or more of the plurality of audio stems having a common source class generated by different source separation processes in the processing recipe, and / or processing one or more of the plurality of audio stems and applying a neural network model trained to clean up artifacts on the audio stems separated according to the processing recipe, thereby mitigating artifacts and / or noise introduced by the sequential audio source separation model.
[0018] At least one source separation process of the method may further comprise the step of separating and outputting the mixture of speech source and a residual complement comprising music and noise.
[0019] The method may further include evaluating output audio stems from a plurality of sequential divergent processing orders to assess an optimized processing order, where the evaluating step includes processing the input mixture to separate a first source class and a first residual complement, then processing the first residual complement, separating a second source class, and generating a first set of audio output stems, and the evaluating step further includes processing the input mixture to separate the second source class and a second residual complement, then processing the second residual complement, separating the first source class, and generating a second set of audio output stems. The method may further include comparing the first set of audio output stems and the second set of audio output stems to determine a processing order that generates higher quality audio output stems.
[0020] The step of separating one or more of the audio sources from the single track audio input sample using the sequential audio source separation model of the method may further include a user-guided and / or self-iterative process configured to progressively separate the sources from the unknown mixture of audio signals by selecting audio enhancements and / or audio enhancement parameters and fine-tuning the sequential audio source separation model to enable matching of the source signals in the unknown mixture of audio signals, and identifying source classes in the unknown mixture of audio signals and identifying a generic separation model for use in a processing recipe.
[0021] The processing recipe of the recipe may further comprise a set of source class and output stems and corresponding source separation models arranged in a hierarchical branching sequence, wherein at least one class of sources is separated into multiple source class separation stems, and / or the processing recipe includes a step of generating a mixture of multiple source class separation stems as a source class output.
[0022] The method may further include training a sequential audio source separation model using the single track audio input sample by providing the single track audio input sample to a generic source separation model, generating a plurality of initial audio stems corresponding to one or more of the plurality of audio sources, and retraining the generic source separation model, at least in part, using one or more of the plurality of initial audio stems, the generic source separation model comprising a plurality of neural network models, each configured to receive a single channel audio input sample and output one or more source-separated audio stems comprising a source class and a remaining complementary signal mixture.
[0023] The training of the sequential audio source separation model of the method may further include training a plurality of neural network models in an order defined by a processing recipe, and / or the processing recipe comprises a hierarchical branching sequence, each branch comprising one or more of the plurality of neural network models. The training of the sequential audio source separation model of the method may further include iteratively training the sequential audio source separation model based, at least in part, on a plurality of audio stems generated from prior iterations of the sequential audio source separation model, evaluating one or more of the plurality of audio stems by determining a metric associated with the source-separated audio stem and comparing the metric to one or more threshold parameters, and / or adding one or more source-separated audio stems to a training dataset to train the sequential audio source separation model based on the evaluating step.
[0024] The training step may further include generating a training data set of artifacts generated by the process recipe to train the one or more neural networks, receiving a source separation output with the artifacts, and generating an enhanced output in which the one or more artifacts have been reduced and / or removed according to the process recipe.
[0025] In various implementations, the system includes a memory component that stores machine-readable instructions and a logic device configured to execute the machine-executable instructions to separate one or more audio sources from a single-track audio input sample comprising an unknown mixture of audio signals generated from the multiple audio sources using a sequential audio source separation model by: defining a processing recipe including a multiple source separation process configured to receive an audio input mixture and output one or more separated source signals and a remaining complementary signal mixture; and processing the single-track audio input sample according to the processing recipe to generate multiple audio stems separated from the unknown mixture of audio signals.
[0026] The logic device may further be configured to define a processing recipe by: processing the single track audio input samples using a first processing order of the multiple source separation processes for separating the audio stems to generate a first set of audio stems; processing the single track audio input samples using a second processing order of the multiple source separation processes that differs from the first processing order for separating the audio stems to generate a second set of audio stems; and evaluating the first set of audio stems and the second set of audio stems to determine which of the first processing order and the second processing order should be included in the processing recipe.
[0027] The logic device may further be configured to define a processing recipe by post-processing the plurality of audio stems to remove and / or mitigate artifacts introduced by the sequential audio source separation model, wherein the post-processing includes processing the plurality of audio stems through one or more neural networks trained to remove and / or mitigate artifacts comprising clicks, harmonic distortion, ghosting, and / or broadband noise.
[0028] The system may be further defined in that at least one source separation process includes separating and outputting a mixture of a speech source and a residual complement comprising music and noise.
[0029] The logic device of the system may further be configured to define the processing by executing multiple source separation processes in a sequential, branching processing order, where a first source separation process is configured to receive a single track audio input sample and output one or more source separation signals and a remaining complementary signal mixture, and where each subsequent separation process is configured to receive an output signal from a previous source separation process, and where each node of the processing recipe is configured to separate a particular source class from the input audio mixture.
[0030] The logic device may be further configured to evaluate the output audio stems from the multiple sequential and divergent processing orders and assess an optimized processing order by: processing the input mixture and separating a first source class and a first residual complement, then processing the first residual complement and separating a second source class, and generating a first set of audio output stems; processing the input mixture and separating a second source class and a second residual complement, then processing the second residual complement and separating the first source class, and generating a second set of audio output stems; and comparing the first set of audio output stems and the second set of audio output stems and determining a processing order that generates higher quality audio output stems.
[0031] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter or to limit the scope of the claimed subject matter. A more extensive presentation of the features, details, utilities, and advantages of the present methods as defined in the claims is provided in the following written description of various implementations of the present disclosure and illustrated in the accompanying drawings.
[0032] Where patent specifications, other external documents, or other sources of information are referenced herein, this is generally for the purpose of providing a context for discussing features of the present invention. Unless specifically stated otherwise, reference to such external documents or such sources of information shall not be construed as an admission that such documents or such sources of information are prior art or form part of the common general knowledge in the art in any jurisdiction. The present invention provides, for example, the following items. (Item 1) 1. A method comprising: receiving single track audio input samples comprising an unknown mixture of audio signals generated from multiple audio sources; defining a processing recipe including a plurality of source separation processes configured to receive an audio input mixture and output one or more separated source signals and a remaining complementary signal mixture; processing the single track audio input samples according to the processing recipe to generate a plurality of audio stems separated from the unknown mixture of audio signals; separating one or more of the audio sources from the single track audio input samples using a sequential audio source separation model comprising: A method comprising: (Item 2) Defining the process recipe comprises: processing the single track audio input samples using a first processing order of the multiple source separation processes for separating audio stems to generate a first set of audio stems; processing the single track audio input samples using a second processing order of the multiple source separation processes that is different from the first processing order for separating audio stems to generate a second set of audio stems; evaluating the first set of audio stems and the second set of audio stems to determine which of the first processing order and the second processing order should be included in the processing recipe; The method according to item 1, comprising: (Item 3) 3. The method of claim 1, wherein the processing recipe further includes post-processing the plurality of audio stems to remove and / or mitigate artifacts introduced by the sequential audio source separation model, the post-processing including processing the plurality of audio stems through one or more neural networks trained to remove and / or mitigate artifacts comprising clicks, harmonic distortion, ghosting, and / or broadband noise. (Item 4) 4. The method of any one of items 1-3, wherein at least one source separation process includes separating and outputting a mixture of a speech source and a residual complement comprising music and noise. (Item 5) the processing recipe further includes executing the plurality of source separation processes in a sequential, branching processing order, wherein a first source separation process is configured to receive the single track audio input samples and output one or more source separation signals and a remaining complementary signal mixture, and each subsequent separation process is configured to receive an output signal from a previous source separation process; each node of said processing recipe is configured to separate a particular source class from an input audio mixture; The method according to item 1. (Item 6) evaluating output audio stems from a plurality of sequential divergent processing orders and assessing an optimized processing order; evaluating includes processing the input mixture to separate a first source class and a first residual complement, then processing the first residual complement to separate a second source class, and generating a first set of audio output stems; evaluating further includes processing the input mixture to separate the second source class and a second residual complement, then processing the second residual complement to separate the first source class, and generating a second set of audio output stems; comparing the first set of audio output stems and the second set of audio output stems to determine a processing order that will produce a higher quality audio output stem; Item 6. The method of item 5, further comprising: (Item 7) The process recipe comprises: combining two or more of the plurality of audio stems having a common source class generated by different source separation processes within the processing recipe; and / or processing one or more of the plurality of audio stems and applying a neural network model trained to clean up artifacts on the audio stems separated according to the processing recipe, thereby mitigating artifacts and / or noise introduced by the sequential audio source separation model. 7. The method of any one of items 1-6, comprising post-processing the plurality of audio stems comprising: (Item 8) Separating one or more of the audio sources from the single track audio input samples using a sequential audio source separation model further comprises: selecting audio enhancements and / or audio enhancement parameters to fine-tune the sequential audio source separation model to enable matching of source signals within the unknown mixture of audio signals; identifying source classes within the unknown mixture of audio signals and identifying a general separation model for use in the processing recipe; 8. The method of any one of items 1-7, comprising a user-guided and / or self-iterative process configured to progressively separate sources from the unknown mixture of audio signals comprising: (Item 9) the processing recipe comprises a set of source classes and output stems and corresponding source separation models arranged in a hierarchical branching sequence; At least one class of sources is separated into a plurality of source class separation stems, and the processing recipe includes generating a mixture of the plurality of source class separation stems as a source class output. The method according to item 8. (Item 10) training the sequential audio source separation model using the single track audio input sample by providing the single track audio input sample to a generic source separation model, generating a plurality of initial audio stems corresponding to one or more of the plurality of audio sources, and retraining the generic source separation model, at least in part, using one or more of the plurality of initial audio stems; the general source separation model comprises a plurality of neural network models, each configured to receive a single channel audio input sample and to output one or more source-separated audio stems comprising a source class and a complementary signal mixture of the remaining samples; 10. The method according to any one of items 1-9. (Item 11) training the sequential audio source separation models further comprises training the plurality of neural network models in an order defined by the processing recipe; the processing recipe comprises a hierarchical branching sequence, each branch comprising one or more of the plurality of neural network models. Item 11. The method according to item 10. (Item 12) 12. The method of claim 11, wherein training the sequential audio source separation model further comprises re-iteratively training the sequential audio source separation model based, at least in part, on the plurality of audio stems generated from a previous iteration of the sequential audio source separation model. (Item 13) Training the sequential audio source separation model further comprises: evaluating one or more of the plurality of audio stems by determining a metric associated with the source-separated audio stem and comparing the metric to one or more threshold parameters; adding one or more source-separated audio stems to a training dataset to train the sequential audio source separation model based on said evaluating; and Item 13. The method according to item 12, comprising: (Item 14) 14. The method of any one of items 10-13, wherein training the sequential audio source separation model further comprises generating a training dataset of artifacts produced by the processing recipe to train one or more neural networks, receiving a source separation output with artifacts, and generating an enhanced output in which one or more artifacts are reduced and / or removed according to the processing recipe. (Item 15) 1. A system comprising: a memory component that stores machine-readable instructions; a logic device, the logic device executing the machine-readable instructions to: defining a processing recipe including a plurality of source separation processes configured to receive an audio input mixture and output one or more separated source signals and a remaining complementary signal mixture; processing the single track audio input samples according to said processing recipe to generate a plurality of audio stems separated from the unknown mixture of audio signals; and separating one or more audio sources from the single track audio input samples comprising an unknown mixture of the audio signals generated from multiple audio sources using a sequential audio source separation model by The logical device and A system comprising: (Item 16) The logic device further comprises: processing the single track audio input samples using a first processing order of the multiple source separation processes for separating audio stems to generate a first set of audio stems; processing the single track audio input samples using a second processing order of the multiple source separation processes that is different from the first processing order for separating audio stems to generate a second set of audio stems; evaluating the first set of audio stems and the second set of audio stems to determine which of the first processing order and the second processing order should be included in the processing recipe; Item 16. The system of item 15, configured to define the process recipe by (Item 17) The logic device further comprises: configured to define a processing recipe by post-processing the plurality of audio stems to remove and / or mitigate artifacts introduced by the sequential audio source separation model; the post-processing includes processing the plurality of audio stems through one or more neural networks trained to remove and / or mitigate artifacts comprising clicks, harmonic distortion, ghosting, and / or broadband noise. Item 17. The system according to item 15 or 16. (Item 18) 18. The system of any one of items 15-17, wherein at least one source separation process includes separating and outputting a mixture of a speech source and a residual complement comprising music and noise. (Item 19) The logic device further comprises: executing the plurality of source separation processes in a sequential, branching processing order, wherein a first source separation process is configured to receive the single track audio input samples and output one or more source separation signals and a remaining complementary signal mixture, and each subsequent separation process is configured to receive an output signal from a previous source separation process; configured to define the process by each node of said processing recipe is configured to separate a particular source class from an input audio mixture; 19. The system according to any one of items 15-18. (Item 20) The logic device further comprises: processing the input mixture to separate a first source class and a first residual complement, then processing the first residual complement to separate a second source class and generate a first set of audio output stems; processing the input mixture to separate the second source class and a second residual complement, then processing the second residual complement to separate the first source class and generate a second set of audio output stems; comparing the first set of audio output stems and the second set of audio output stems to determine a processing order that will produce a higher quality audio output stem; 20. The system of claim 19, configured to evaluate output audio stems from a plurality of sequential divergent processing orders and assess an optimized processing order by: [Brief explanation of the drawings]
[0033] Aspects of the present disclosure and their advantages may be better understood by reference to the following drawings and the following detailed description. It should be understood that like reference numerals are used to identify like elements shown in one or more of the figures, and that the illustrations therein are intended to illustrate implementations of the present disclosure, and not to limit it. The components in the drawings are not necessarily to scale, emphasis instead being placed upon clearly illustrating the principles of the present disclosure.
[0034] [Figure 1] FIG. 1 illustrates an audio source separation system and process, according to one or more implementations.
[0035] [Figure 2] FIG. 2 illustrates elements associated with the system and process of FIG. 1, according to one or more implementations.
[0036] [Figure 3] FIG. 3 is a diagram illustrating a machine learning dataset and training data loader, according to one or more implementations.
[0037] [Figure 4] FIG. 4 illustrates an exemplary machine learning training system that includes a self-iterating dataset generation loop, according to one or more implementations.
[0038] [Figure 5] FIG. 5 illustrates an example operation of a data loader for use in training a machine learning system, according to one or more implementations.
[0039] [Figure 6] FIG. 6 illustrates an exemplary machine learning training method, according to one or more implementations.
[0040] [Figure 7] FIG. 7 illustrates an exemplary machine learning training method, including a training mixture example, according to one or more implementations.
[0041] [Figure 8] FIG. 8 illustrates an exemplary machine learning process, according to one or more implementations.
[0042] [Figure 9] FIG. 9 illustrates an exemplary post-processing model configured to clean up artifacts introduced by machine learning processes, according to one or more implementations.
[0043] [Figure 10] FIG. 10 illustrates an exemplary user-guided self-iterative training loop, according to one or more implementations.
[0044] [Figure 11A] FIG. 11 comprises FIGS. 11A and 11B, which illustrate an exemplary machine learning application, according to one or more implementations. [Figure 11B] FIG. 11 comprises FIGS. 11A and 11B, which illustrate an exemplary machine learning application, according to one or more implementations.
[0045] [Figure 12]FIG. 12 illustrates an exemplary machine learning processing application, according to one or more implementations.
[0046] [Figure 13A] FIG. 13 comprises FIGS. 13A, 13B, 13C, 13D, and 13E, which illustrate one or more examples of a multi-model recipe processing system, according to one or more implementations. [Figure 13B] FIG. 13 comprises FIGS. 13A, 13B, 13C, 13D, and 13E, which illustrate one or more examples of a multi-model recipe processing system, according to one or more implementations. [Figure 13C] FIG. 13 comprises FIGS. 13A, 13B, 13C, 13D, and 13E, which illustrate one or more examples of a multi-model recipe processing system, according to one or more implementations. [Figure 13D] FIG. 13 comprises FIGS. 13A, 13B, 13C, 13D, and 13E, which illustrate one or more examples of a multi-model recipe processing system, according to one or more implementations. [Figure 13E] FIG. 13 comprises FIGS. 13A, 13B, 13C, 13D, and 13E, which illustrate one or more examples of a multi-model recipe processing system, according to one or more implementations.
[0047] [Figure 14] FIG. 14 illustrates an exemplary audio processing system, according to one or more implementations.
[0048] [Figure 15] FIG. 15 illustrates an example neural network that may be used in one or more of the implementations of FIGS. 1-14, according to one or more implementations. DETAILED DESCRIPTION OF THE INVENTION
[0049] Detailed Description In the following description, various implementations will be described. For purposes of explanation, specific configurations and details are set forth to provide a thorough understanding of the implementations. However, it will also be apparent to those skilled in the art that the implementations may be practiced without the specific details. Additionally, well-known features may be omitted or simplified to avoid obscuring the described implementations.
[0050] Improved audio source separation systems and methods are disclosed herein. In various implementations, a single-track (e.g., undifferentiated) audio recording is provided to an audio source separation system configured to separate and separate various audio components, such as speech and instruments, into high-fidelity stems, i.e., discrete or grouped collections of co-mixed audio sources. In various implementations, the single-track audio recording contains an unidentified audio mixture (e.g., the audio sources, recording environment, and / or other aspects of the audio mixture are unknown to the audio source separation system), and the audio source separation system and method are adapted to identify and / or separate the audio sources from the unidentified audio mixture in a self-iterative training and fine-tuning process.
[0051] The systems and methods disclosed herein may be implemented on at least one computer-readable medium carrying instructions that, when executed by at least one processor, cause the at least one processor to perform any of the method steps disclosed herein. Some implementations relate to a computer system including at least one processor and a memory that stores instructions that, when executed by the at least one processor, cause the at least one processor to perform any of the method steps disclosed herein. In various implementations, the models described herein may be implemented as stored data and software modules and / or code that operates on the stored data.
[0052] In some cases, training a model involves processing a data structure or structures to form a new data structure that a component of a computer system can access and use as a model. For example, an artificial intelligence system may include a computer with one or more processors, a program code memory, a writable data memory, and several inputs / outputs. The writable data memory may hold several data structures corresponding to trained or untrained models. Such data structures may represent one or more layers of nodes of a neural network and links between nodes in different layers, and weights for at least some of the links between nodes. In other cases, different types of data structures may represent the model.
[0053] In some cases, when referring to training a model, feeding a model, and / or having a model take in inputs and provide outputs, this may refer to the action of a computer capable of reading a writable data memory that contains the model and executes program code for working with the model. For example, the model may be trained with a set of training data, which may be the training examples themselves and / or training examples and corresponding ground truth. Once trained, the model may be usable to make decisions about examples provided to the model. This may be done by the computer receiving, at input, input data representing the examples, performing a process using the examples and the model, and outputting, at output, output data representing and / or indicating decisions made by or based on the model.
[0054] In a very specific example, an artificial intelligence system may have a processor that reads in multiple photos of cars and reads in ground truth data indicating "these are cars." The processor may read in multiple photos of street lamps and the like and reads in ground truth data indicating "these are not cars." In some cases, the model is trained on the input data itself, without being provided with ground truth data. The result of such processing may be a trained model. The artificial intelligence system can then, with the model trained, be provided with an image without any indication of whether it is an image of a car, and can output an indication of a determination of whether it is an image of a car.
[0055] To process audio signals, data, recordings, etc., the input may or may not be the audio data itself and some ground truth data about the audio data. Then, once trained, the artificial intelligence system may receive some unknown audio data and output decision data about the unknown audio data. For example, the output decision data may relate to extracted notes, stems, frequencies, etc., or other decisions or AI-determined observations of the input audio data.
[0056] The resulting data structure corresponding to the trained model can then be ported or distributed to other computer systems, which can then use the trained model. Once trained, the computer code in program memory, when executed by a processor, can receive an image at input, and based on the fact that the data structure represents the trained AI model being trained, the program code can process the input and output a decision regarding the nature of the input. In some implementations, an AI model may comprise a data structure and program code that are intertwined and not easily separated.
[0057] A model may be represented by a set of weights assigned to edges in a neural network graph, program code and data containing instructions regarding the neural network graph and how to interact with the graph, mathematical expressions such as a regression model or classification model, and / or other data structures as may be known in the art. A neural model (or neural network, or neural network model) may be embodied as a data structure that shows or represents a set of connected nodes, often referred to as neurons, many of which may be data structures that mimic or simulate the signal processing performed by biological neurons. Training may include updating parameters associated with each neuron, such as the other neurons to which it is connected and the weights and / or functions of the neuron's input and output. In practice, a neural model may pass and process input data variables in some way to generate output variables to achieve some goal, for example, to generate a binary classification of whether an input image or input dataset fits a certain category. The training process may involve complex calculations (e.g., calculating gradient updates and then using the gradients to update parameters layer by layer). The training may be performed using some parallel processing.
[0058] Any of the models for audio source separation described herein may also be referred to, at least in part, by several terms including, for example, an “audio source separation model,” “recurrent neural network,” “RNN,” “deep neural network,” “DNN,” “inference model,” or “neural network.”
[0059] 1 and 2 illustrate an audio source separation system and process according to one or more implementations of the present disclosure. The audio processing system 100 includes a core machine learning system 110, a core model operation 130, and a modified recurrent neural network (RNN) class model 160. In the illustrated implementation, the core machine learning system 110 implements an RNN class audio source separation model (RNN-CASSM) 112, depicted in a simplified representation in FIG. 1. As shown, the RNN-CASSM 112 receives a signal input 114, which is input to a time-domain encoder 116. The signal input 114 includes a single-channel audio mixture, which may be received from a stored audio file accessed by the RNN-CASSM network, an audio input stream received from a separate system component, or another audio data source. The time-domain encoder 116 models the input audio signal in the time domain and estimates audio mixture weights. In some implementations, the time-domain encoder 116 segments the audio signal input into separate waveform segments that are normalized for input to a 1D convolutional coder. The RNN class mask network 118 is configured to estimate source masks for separating audio sources from the audio input mix. The masks are applied to the audio segments by a source separation component 120. The time-domain decoder 122 is configured to reconstruct the audio sources, which are then available for output through a signal output 124.
[0060] In an exemplary implementation, the RNN-CASSM 112 is modified for operation with the core model operations 130, as will be described herein. Referring to block 1.1, audio source data is sampled using a 48 kHz sample rate in various implementations. Thus, the RNN-CASSM 112 may be trained at higher than audible frequencies to recognize distinct stems of audio in lower frequency ranges. It has been observed that speech separation model implementations trained at sample rates lower than 48 kHz produce lower quality audio separation in audio samples recorded with older equipment (e.g., Nagra™ equipment from the 1960s). Various steps disclosed herein, such as training the source separation model, setting appropriate hyperparameters for that sample rate, and operating the signal processing pipeline at the 48 kHz sample rate, are performed at the 48 kHz sample rate. Other sampling rates, including oversampling, may also be used in other implementations, consistent with the teachings of this disclosure. In block 1.2, the encoder / decoder framework (eg, the time domain encoder 116 and the time domain decoder 122) is set to a step size of one sample (eg, the input signal sample rate).
[0061] Referring to block 1.3, a modified RNN-CASSM 160 is generated that extends beyond the source separation and noise reduction of conventional implementations. It is observed that processing audio mixtures with a trained RNN-CASSM can sometimes result in undesirable click artifacts, harmonic artifacts, and broadband noise artifacts. To address this, the audio processing system 100 applies modifications to the RNN-CASSM network 112, as referenced in blocks 1.3a, 1.3b, and / or 1.3c, to reduce and / or avoid the artifacts and provide other benefits. The same modifications may also be used for more transformative types of trained processing as well.
[0062] The model operation associated with the modified RNN-CASSM 160 will now be described in further detail with reference to blocks 1.3ac. In block 1.3a, at least some of the encoder and decoder layers are removed and not used in the modified RNN-CASSM. It is observed that these removed layers may be redundant in the trained filter when the step size is one sample. In block 1.3b, the step of applying a mask (e.g., component 120) is also removed for at least some of the audio processing. Thus, the modified RNN-CASSM 160 may perform audio source separation without applying a mask. The masking step potentially contributes to click artifacts often present in the generated audio stem. To address this, the audio processing system may omit some or all of the masking and instead use the output of the RNN class mask network 118 in a more direct manner (e.g., training the RNN to output one or more separated audio sources). In block 1.3c, a window function is applied to the overlap-add step of an RNN class mask network (e.g., RNN network 162). When the model output has linear harmonic series "banding" artifacts at frequencies related to the overlap-add function segment length, the audio processing system can extend the window function across each overlapping segment to smooth hard edges when reconstructing the audio signal.
[0063] Referring to block 1.4, another core model operation 130 is the use of a separation strength parameter, which allows control over the strength of the separation mask that is applied to the input signal to generate the separated sources. To provide direct control over the strength of the separation mask that is applied to the input signal to generate the separated sources, a parameter that determines how strongly the separation mask is applied is introduced during the forward pass of the model. In one embodiment, the separation strength parameter is given by f(M)=M Sor the like, where the mask M has values [0, l] and s is a separation strength parameter. In this example, values of s > 1.0 result in lower mask values and fewer target sources when the mask is applied to the input mixture, and values of s < 1.0 result in higher mask values and a combination of target sources and complementary and noise components in the signal.
[0064] The separation strength parameter may be implemented as a helper function for an automated version of the Self-Iterative Process Training (SIPT) algorithm, which will be described in further detail below. It should be understood that the RNN-CASSM 112 may implement one or more of the core model operations 130 disclosed herein and may include additional operations consistent with the teachings of this disclosure.
[0065] 3-5, an exemplary process for training an RNN-CASSM network for audio source separation will now be described. The training process includes a plurality of labeled machine learning training datasets 310 configured to train a network for audio source separation as described herein. For example, in some implementations, the training dataset may include an audio mix and ground truth labels that identify source classes to be separated from the audio mix. In other implementations, the training dataset may include separated audio stems having audio artifacts (e.g., clicks) generated by the source separation process and / or one or more audio enhancements (e.g., reverb, filters), and ground truth labels that identify enhanced audio stems in which the identified audio artifacts and / or enhancements have been removed.
[0066] In operation, the network is trained by feeding labeled audio samples into the network. In various implementations, the network includes multiple neural network models that can be separately trained for specific source separation tasks (e.g., separating speech, separating foreground speech, separating drums, removing artifacts, etc.). Training involves a forward pass through the network to generate audio source separation data. Each audio sample is labeled with a "ground truth" that defines the expected output, which is compared to the generated audio source separation data. If the network mislabels an input audio sample, a backward pass through the network may be used to adjust the network's parameters to correct the misclassification. In various implementations, [ka] The output estimate, denoted as Y, is compared to the ground truth, denoted as Y, using a regression loss, such as an L1 loss function (e.g., least absolute deviation), an L2 loss function (e.g., least squares), a scale-invariant signal-to-distortion ratio (SISDR), a scale-dependent signal-to-distortion ratio (SDSDR), and / or other loss functions as known in the art. After the network is trained, a validation dataset (e.g., a set of labeled audio samples not used in the training process) may then be used to measure the accuracy of the trained network. The trained RNN-CASSM network may then be implemented in a runtime environment to generate distinct audio source signals from the audio input stream. The generated distinct audio source signals may also be referred to as generated multiple audio stems. The generated multiple audio stems may correspond to one or more audio sources of the multiple audio sources of the audio input stream.
[0067] In various implementations, the modified RNN-CASSM network 160 is trained based on multiple (e.g., thousands) audio samples, including audio samples representing multiple speakers, instruments, and other audio source information under various conditions (e.g., including various noisy conditions and audio mixes). From the error between the separated audio source signals and labeled ground truth associated with audio samples from the training dataset, the deep learning model learns parameters that enable the model to separate the audio source signals. The RNN-CASM 120 and the modified RNN-CASM 160 may also be referred to as the inference model 120 and the modified inference model 160 and / or the trained audio separation model 120 and the updated audio separation model 160.
[0068] 3 is a schematic diagram of a machine learning dataset 310 and training data loader 350 as may be relevant to the dataset and dataset manipulation during training by an audio processor that may provide improved and / or useful functionality in output signal quality source separation. In the illustrated implementation, an RNN-CASSM network is trained to separate audio sources from a single-track recording of a music recording session that includes a mixture of people singing / talking, instruments being played, and various ambient noises.
[0069] Referring to block 2.1, the training dataset used in the illustrated implementation includes a 48 kHz speech dataset. For example, the 48 kHz speech dataset may include identical speech recorded simultaneously at various microphone distances (e.g., a close microphone and a more distant microphone). In one test implementation, 85 different speakers were included in the 48 kHz speech dataset, with more than 20 minutes of speech per speaker. In various implementations, an exemplary speech dataset may be created using a large number of speakers, such as 10, 50, 85, or more, from adult male and female speakers, with extended speech periods, such as 10, 20, or more minutes, and recordings at 48 kHz or higher sampling rates. It should be understood that other dataset parameters may also be used to generate a training dataset in accordance with the teachings of the present disclosure.
[0070] Referring to block 2.2, the training dataset further includes a non-speech music and noise dataset that includes segments of input audio, such as digitized mono audio recordings originally recorded on analog media. In some implementations, this dataset may include segments of recorded music, non-vocal sounds, background noise, audio media artifacts from digitized audio recordings, and other audio data. Using this dataset, an audio processing system can more easily separate the voice of a speaker of interest from other voices, music, and background noise in the digitized recording. In some implementations, this may include using manually collected segments of the recording that are manually annotated as lacking speech or another audio source class and labeling those segments accordingly.
[0071] Referring to block 2.3, a dataset is generated and refined using an incremental, self-iterative dataset generation process using a target unidentified mixture (e.g., a mixture unknown to the audio processing system). The generated dataset may include a labeled dataset generated by processing an unlabeled dataset (e.g., a target unidentified mixture to be source-separated) through one or more neural network models to generate an initial classification. This "coarsely separated data" is then processed through a pruning step configured to select from among the "coarsely separated data" to retain the most useful "coarsely separated data" based on a usefulness metric. For example, the performance of the training dataset may be measured by applying a validation dataset to a model trained using various training datasets, including the "coarsely separated data," and determining, based on the calculated validation error, data samples that contribute to better and poorer performance. The usefulness metric may be implemented as a function that estimates a quality metric in the coarsely separated data to identify "low-quality" fine-tuning data that should be discarded before the next fine-tuning iteration. For example, a moving root-mean-square (RMS) window function may be calculated on the network's output to identify segments of the output where the RMS metric lies above a calibrated threshold for a certain (or minimum) duration of samples. This metric can be used, for example, to identify low-amplitude segments of coarsely separated source outputs where artifacts are more likely to occur. The threshold and minimum duration may be user-adjustable parameters, allowing for adjustment of the data to be discarded.
[0072] The progressively self-iterative dataset generated using the target unidentified mixture may be generated using a self-iterative dataset generation loop 420 as illustrated in FIG. 4. In various implementations, previous recordings containing sources to be separated that are not yet identified by the trained source separation model are less likely to be successfully separated by the model. The existing dataset may not be substantial enough to train a robust separation model for the targeted sources, and opportunities to capture new recordings of the sources may not exist. Instead of capturing new recordings of the sources, additional training data may be manually labeled from isolated instances of the sources in previous recordings. For example, isolated speech from an identified speaker, isolated audio of an identified instrument recorded with similar equipment in a similar environment, and / or other available audio segments that approximate the source being separated may be added to the training data manually and / or automatically (e.g., based on metadata labeling of the audio source, such as source identification, source class, and / or environment). This additional training data may be used to help fine-tune the model to improve processing performance on previous recordings. However, this labeling process can involve a significant amount of time and manual effort, and there may not be enough isolated instances of the source in previous recordings to provide a sufficient amount of additional training data.
[0073] The illustrated generative refinement tool can overcome these difficulties. In one method, a coarse generic model 410 is trained on a generic training dataset. The generic training dataset may include labeled source audio data and labeled noise audio data. The generic model 410 may be referred to as a generic source separation model 410 or a trained audio source separation model 410. The training dataset may include multiple datasets, each of which may include labeled audio samples configured to train the system to address the source separation problem. The multiple datasets may include a speech training dataset including multiple labeled speech samples and / or a non-speech training dataset including multiple labeled music and / or noise data samples. Available previous recordings containing the unidentified audio mixture to be separated are then processed with the generic model in process 422, resulting in two labeled datasets of isolated audio (e.g., audio stems): a coarse separated previous recording source dataset 424 and a coarse separated previous recording noise dataset 426.
[0074] In various implementations, other training datasets may also be used, providing a set of labeled audio samples selected to train the system to solve a particular problem (e.g., speech vs. non-speech). In some implementations, for example, the training dataset may include (i) music vs. sound effects vs. foley, (ii) a dataset for various instruments in a band, (iii) multiple human speakers separated from each other, (iv) sources from room reverberation, and / or (v) other training datasets. The results are then culled (process 428) using a threshold metric to remove audio windows that fall below a selected root-mean-square (RMS) level, which may be user-selectable. In some implementations, a running RMS may be calculated by segmenting the audio data into overlapping windows of equal duration and calculating the RMS for each window. The RMS level may be referred to as a usability metric or a quality metric, and the RMS level may be one option for alternative usability or quality metrics. The quality metric may be calculated based on multiple associated audio stems.
[0075] A new model is then trained in process 430 using the culled self-repeated dataset (e.g., the culled result added to the audio training dataset, the culled self-repeated dataset is also referred to as a culled dynamically evolving dataset) to train and generate an improved model 432 and improve its performance when processing the recording. The improved model 432 may be configured to reprocess the audio input stream and generate multiple improved audio stems. This improved model is an update of the trained audio source separation model. This improved model may be referred to as an updated audio source separation model 432.
[0076] In some implementations, the audio training dataset may be curated during an iterative fine-tuning process to remove data that is not relevant to the target input mixture. For example, the input mixture may identify / classify various sources within the input mixture, leaving certain other source categories that are not identified / classified and / or otherwise not relevant to the source separation task. Training data associated with these "irrelevant" source categories (e.g., categories not found in the target mixture, categories identified by the user as not relevant to the source separation task, and / or other irrelevant source categories as defined by other criteria) may be culled from the audio training dataset, allowing the training dataset to become increasingly specific to the content of the target input mixture.
[0077] This process 420 is repeated iteratively, each time improving the separation quality of the model (e.g., fine-tuning to improve the accuracy and / or quality of source separation). In response to subsequent iterations, process 420 may use additional RMS levels, whereby the additional RMS levels exceed the previous RMS levels. This process allows for more automated refinement of the initial general source isolation or separation model. It has been observed that looping at various stages shows improvement over larger related mixtures. The general model 410 and the improved model 432 may also be referred to as the inferred model 410 and the modified inferred model 432 and / or the trained audio separation model 410 and the updated audio separation model 432.
[0078] Improvements in separation quality (e.g., audio fidelity) can be measured by the system and / or evaluated by a user providing feedback to the system through a user interface and / or overseeing one or more steps in process 420. In some implementations, process 420 may use a combination of an algorithm and / or user ratings to calculate a mean opinion score (MOS) to estimate separation quality. For example, an algorithm can estimate the amount of artifacting generated during the source separation operation, which in turn relates to the overall quality of the network's output. In some implementations, estimating the amount of artifacting generated during the source separation operation includes feeding the separated sources through a neural networking model trained to separate audio artifacts from signals, allowing for measurement of the presence and / or strength of such audio artifacts. The strength of the audio artifacts can be determined at each iteration and tracked between iterations to fine-tune the model. In some implementations, the iterative process continues until the estimated separation quality across iterations no longer improves and / or the estimated separation quality meets one or more predetermined quality thresholds.
[0079] Referring to blocks 2.4 and 2.4a, the machine learning training data loader 350 is configured to match the sound qualities (e.g., perceived distance from the microphone, filtering, reverberation, echo, nonlinear distortion, spectral distribution, and / or other measurable audio qualities) of the target unidentified mixture during training. A challenge with training an effective supervised source separation model is that the dataset example should ideally be curated to match as closely as possible the qualities of the sources in the target mixture. For example, if a speaker should be isolated from the mixture in which they are speaking into a microphone in a reverberant hall, and the recording is captured from some distance with a recording device in the audience, a comparison may be made with what the same speaker might sound like when recorded speaking directly into a microphone in a neutral, non-reverberant space, such as where a high-quality speaker dataset might have been recorded. The goal then becomes adding enhancements to the high-quality speech dataset samples during training that generally result in lower deviations from the target input mixture. In this example, a "reverberation" extension may be added to simulate that of a hall, a "nonlinear distortion" extension to simulate the voice of a speaker amplified by a sound system, and a "filter" extension to simulate the distance of the sound system from the recording device.
[0080] In various implementations, the solution involves a hierarchical mix-bus schema, including pre- and post-mixture extension modules. Creating an ideal target mixture would be tedious or impractical to create manually each time a new type of mixture is required for training. The hierarchical mix-bus schema allows for easy definition of arbitrarily complex randomized "sources" and "noise" during supervised source separation training. The data loader roughly matches qualities such as filtering, reverberation, relative signal level training, and other audio qualities for improved source separation enhancement results. The machine learning data loader uses a hierarchical schema that allows for easy definition of "source" and "noise" mixtures that are dynamically generated from source data while training the model. The mix-bus allows for optional extensions such as reverb or filters with accompanying randomization parameters. Using appropriately classified dataset media as raw materials, this allows for easy creation of training datasets that mimic the desired source separation target mixture.
[0081] An exemplary simplified schema representation 550 is illustrated in Figure 5. The training mixture schema includes separate options for sources and noise, including criteria such as dB ranges, probabilities associated with source determination, room impulse responses, filters, and other criteria.
[0082] Referring to block 2.4b, the data loader also provides pre- and post-mixture augmentations, including filters, nonlinear functions, and convolutions, that are applied to the target mixture during training. In various implementations, relevant augmentations are identified and added to the training dataset, and the separated audio stems are post-processed (e.g., using a pipeline described herein). The source separation model can be trained with an additional goal of transforming the separated sources using the augmentations. In some implementations, the system can be trained to strictly isolate sources as they can be heard in the input mixture. In some implementations, the system may further be trained to improve the quality of some of the separated sources by applying appropriate augmentations, e.g., augmentations that result in minimal deviation from the target mixture when using the available training dataset. The deviations and appropriate augmentations may be estimated algorithmically and / or by user evaluation. For example, a voice in the input mixture may be filtered and difficult to understand due to being recorded behind a closed door. In this example, the separated sources may be augmented (e.g., to degrade the input audio dataset and generally match the sound of a voice behind a closed door). However, the target separated source outputs during training are not augmented in this example (e.g., augmented audio inputs during training vs. a corresponding high-quality audio target output dataset), and thus the network is trained to approximate this same transformation by post-processing augmentation.
[0083] In some implementations, the target source may be a transformed version of its current representation in the mixture. For example, there may be a need to restore a bandwidth-limited recording to a more complete frequency spectrum, or to isolate and increase the proximity fidelity of an obscured background speaker. These needs may be user-determined or left to the transformative model itself to resolve automatically based on input deviations from the target output training set. For example, if the model is trained to output high-quality near-neighbor speech without much reverberation when using a randomly expanded speech dataset during training, inputting mixtures containing such high-quality near-neighbor speech without much reverberation may tend to result in minimal changes to those inputs. However, inputting mixtures containing speech that deviates from these qualities may tend to transform those input speech mixtures to resemble high-quality near-neighbor speech without much reverberation.
[0084] The augmentation module can be used by the data loader to generate training examples consisting of an augmented source as input, along with alternatively augmented versions of the same source as the target output. This enables transformational examples in which the target source, when part of the training mixture, can be represented in an alternatively augmented context. When used while training a modified RNN-CASSM, this allows the audio processing system to learn operations such as "filtering out," "dereverberating," and deeper restoration of highly obscured target sources.
[0085] Application examples 500 in the illustrated implementation include (i) filtering removal, which includes expanding the mixture using a filter; (ii) dereverberation, which includes expanding the mixture using a reverb; (iii) background speaker recovery, which includes expanding the source using a filter and a reverb; (iv) distortion repair, which includes expanding the mixture using a distortion; and (v) gap repair, which includes expanding the mixture using a gap.
[0086] 6 and 7, an exemplary implementation of a machine learning training method 600 will now be described. In these examples, the machine learning training method will be described in relation to a method of training for improvements and / or useful functionality in output signal quality and source separation. Referring to block 3.1, a first machine learning training method includes upscaling the trained network sample rate (e.g., from 24 kHz to 48 kHz). Due to limitations in time, computational resources, etc., the model is trained at 24 kHz using a 24 kHz dataset, but this may entail limitations on output quality. An upscaling process undertaken on a model trained at 24 kHz can provide functionality at 48 kHz. One exemplary process includes preserving the inner blocks of learned parameters of the masking network while discarding the encoder / decoder layers and their direct connections. In other words, only the inner separation layers are transplanted into a newly initialized model with a 48 kHz encoder / decoder. Next, the untrained 48 kHz encoder / decoder and its direct connections are fine-tuned using the 48 kHz dataset while the inherited network remains frozen. This is done until an acceptable validation / loss value (e.g., L1, L2, SISDR, SDSDR, or other loss calculation) is again identified during training / validation, which here refers to the adapted inherited layers. For example, the acceptable validation / loss value may be determined by observing a trend toward values identified in previous training sessions compared to model performance, by comparison to a predetermined threshold loss value, or by other approaches. During training, the loss value ideally tends toward minimization; however, in practice, the loss value can also be useful in signaling significant problems during training, such as when the loss value begins to trend away from minimization. It has also been observed that while the loss value may not have improved during training, the model's performance, as measured by the quality of source separation, can still improve by continuing training.
[0087] Finally, fine-tuning training continues across all layers, allowing the model to further evolve at 48 Khz, with the end result being a well-performing 48 Khz model. In some implementations, the system is trained to operate more quickly at high signal processing sample rates by training at a lower sample rate and inheriting appropriate layers into the higher sample rate model architecture, performing a two-step training process: first, untrained layers are trained while the parameters of the inherited layers are frozen; and second, the entire model is then fine-tuned until performance meets or exceeds that of the lower sample rate model. This process can be performed over many iterations, including, but not limited to, the following:
[0088] a) Train at 6 kHz
[0089] b) Upscaling to 12kHz
[0090] c) Upscaling to 24kHz
[0091] d) Upscaling to 48kHz
[0092] Referring to block 3.2, multiple audio source mixtures are used to improve the performance of the speech isolation model (e.g., source = foreground, background, and distant speech, noise, and music mixtures). Speech initially trained on a single speech-to-noise / music mixture may not perform well; the processed results may have difficulty consistently extracting from the original source media and suffer from substantial artifacts. Instead of training with a single speech-to-noise mixture, the audio processing system may provide substantially improved results by using layered multiple audio sources to simulate variations in proximity, such as foreground and background speech. This approach can also be applied to musical instruments, e.g., multiple overlaid guitar samples within a mixture, instead of only one sample at a time. In various implementations, training samples with various layer scenarios may be selected to match and / or approximate the unidentified audio mixture (e.g., based on user input, identified source classes, and / or analysis of the unidentified audio mixture during iterative training).
[0093] The example training mixture 700 includes a speech isolation training mixture 702 that includes a mixture of sources and noise 704. The source mixture 706 may include a mixture of foreground speech 708, background speech 710, and distant speech 712 in an expanded, randomized combination. The noise mixture 714 includes a mixture of instruments 716, room tones 718, and hiss 720 in an expanded, randomized combination.
[0094] 8-10, an implementation of machine learning process 800 will now be described in connection with a method of processing using a machine learning model that contributes to improvements and / or useful functionality in output signal quality and source separation (e.g., as discussed previously herein). Referring to block 4.1, the machine learning process may include imputing the sum of the separated sources as an additional output. In other implementations, the model outputs speech and discarded music / noise. The imputed output may subsequently be used in various processes to further process / separate the sources remaining in the imputed output.
[0095] Referring to block 4.2, the machine learning post-processing model cleans up artifacts introduced by the machine learning process, such as clicks, harmonic distortion, ghosting, and broadband noise (artifacts may be determined, for example, as discussed above with reference to FIG. 4). The trained source separation model may exhibit artifacts such as clicks, harmonic distortion, broadband noise, and "ghosting," where sounds are partially separated between the target complement outputs. In order for these outputs to be used in the context of a high-quality soundtrack, laborious cleanup would typically need to be attempted using conventional audio restoration software. Such attempts may still result in undesirable quality in the restored audio. This can be addressed by post-processing the processed audio with a model that has been trained on a dataset consisting of processing artifacts. The processing artifact dataset may be generated by the problematic model itself.
[0096] Once trained, the post-processing model 910 can be reused for all similar models. In the illustrated implementation, an input mixture 950 is processed using a generic model 952, which generates a source separation output with machine learning artifacts (step 954). A post-processing step 956 removes the artifacts and generates an enhanced output 960. The post-processing model 910 includes generating a dataset 914 of isolated machine learning artifacts (step 912). The machine artifacts may include clicks, ghosting, broadband noise, harmonic distortion, and other artifacts. The isolated machine learning artifacts 912 are used to train a model that removes the artifacts in step 916.
[0097] Referring to block 4.3, in some implementations, a user-guided self-iterative processing / training approach may be used. The user guides and contributes to the fine-tuning of the pre-trained model over the course of the processing / editing / training loop, which can then be used to produce better source separation results than would have been possible from a general model. Model fine-tuning capabilities can be placed in the hands of the user to solve source separations that the pre-trained model may not be able to solve. Processing with a source separation model on an unidentified mixture is not always successful, usually due to a lack of sufficient training data. In one solution, the audio processing system uses a method whereby the user can guide and contribute to the fine-tuning of the pre-trained input, which can then be used to produce better results. In an exemplary method, i) the user processes the input media; ii) is given the opportunity to assess the output, or has the option to have this assessment performed by an algorithm that measures threshold parameters for some metric (e.g., measured using a moving RMS window and / or other measurements as discussed above); iii) if the output is deemed acceptable, processing ends here; otherwise, iv) the user is given the opportunity to manipulate the imperfect output using temporal and / or spectral editing and / or culling / expansion algorithms. Essentially, a segmentation of the output that will best serve the immediate step is selected. In step v), the media is now considered for inclusion in the training dataset; in step vi), the user is also given the opportunity to add their own auxiliary dataset; in step vii), the model is fine-tuned and trained according to the user's hyperparameter preferences; in step viii), the model's performance is validated to confirm improved results from previous iterations; and then in step ix), the process is repeated. In various implementations, hyperparameters associated with fine-tuning training may include parameters such as training segment length, epoch duration, training scheduler type and parameters, optimizer type and parameters, and / or other hyperparameters.The hyperparameters may be initially based on a set of predetermined values and then modified by the user for fine-tuning training.
[0098] An exemplary user-guided self-iteration process 1000 is illustrated in FIG. 10 . Process 1000 begins with a generic model 1002, which may be implemented as a pre-trained model as discussed above. An input mixture 1004, such as a single-track audio signal with an unidentified source mixture or multiple single-track audio signals, is processed through generic model 1002 in step 1006 to generate separated audio signals from the mixture, including a machine-learning separated source signal 1008 and a machine-learning separated noise signal 1010. In step 1012, the results are evaluated to ensure the separated audio sources are of sufficient quality (e.g., by comparing estimated MOS and thresholds and / or other quality measures as described above). If the results are determined to be good, the separated sources are output in step 1014.
[0099] If it is determined that the separated audio source signals require further improvement, one or more of the outputs are prepared for inclusion in a training dataset for fine-tuning in step 1016. In the automated fine-tuning system 1018, the machine learning separated sources 1008 and noise 1010 are used directly as a fine-tuning dataset 1034 (step 1022). The fine-tuning dataset is optionally culled in step 1036 based on a user-selected threshold of a source separation metric (e.g., comparing a moving RMS window to a threshold and / or other separation metrics as described above). Training is then performed in step 1038 to fine-tune the model. The fine-tuned model 1032 is then applied to the input mixture in step 1006.
[0100] In the user-guided fine-tuning system 1020, the user may select portions of audio clips to include or omit from the fine-tuning dataset (step 1024—temporal editing). The user may also select portions of clips to include / omit from the fine-tuning dataset and frequency / time window selections (step 1026—spectral editing). In some implementations, the user may provide additional audio clips to extend the fine-tuning dataset (step 1028—add to dataset). In some implementations, the user provides equalization, reverb, distortion, and / or other enhancement settings to fine-tune the dataset (step 1030—expand). After the fine-tuning dataset 1034 is updated, the training process continues with steps 1036-1038, generating a fine-tuned model 1032.
[0101] Referring to block 4.4, an animated visual representation of the model fine-tuning progress may be implemented. While a user fine-tunes a model to solve source separation for a particular media clip, the model's progressive output is displayed to help guide the user's decision-making and to help indicate the model's performance. In some implementations, for example, an interface may be displayed in a window with associated tool icons that displays a periodically updated spectrogram animated representation of the estimated outputs as computed by the fine-tuned model, which is periodically tested while it is being fine-tuned. This interface can allow the user to visually assess how well the model is currently performing in various regions of the input mixture. The interface may also facilitate user interaction, such as allowing the user to experiment with time / frequency selection of these estimated outputs based on interaction with the spectrogram window.
[0102] Referring to block 4.5, user-guided extensions for fine-tuning training may be implemented. To improve results when enhancing / separating targeted recordings with specific characteristics such as reverberation, filtering, nonlinear distortion, etc., the audio processing system may present the user with tools to control the underlying algorithms for guiding and contributing to the selection of extensions, such as reverberation, filtering, nonlinear distortion, and / or noise, during step 1020 of the loop described in block 4.3. The user can guide and contribute to the selection of extensions and / or extension parameters used during fine-tuning of the pre-trained model over the course of the processing / editing / training loop. The extensions include user-controllable randomization settings for each parameter (e.g., values affecting various aspects of the extension, such as the intensity, density, modulation, and / or decay of the reverberation extension) to help generalize or narrow the targeted behavior after fine-tuning. This allows for greater control over the fine-tuning training, allowing them to be specifically matched to, for example, the filtered sounds / reverberations in the targeted recording. When referring to a random process or randomization, it may be sufficient to have a pseudo-random process or an arbitrary selection process. In some implementations, the expansion parameters are automatically matched to the input source mixture using one or more algorithms. In some implementations, the automatic match is used to achieve a rough match as a starting point. For example, a spectral analysis of the target input mixture may be combined with an analysis of the random dataset samples to yield a set of frequency band deviation scores that can be used to minimize deviation between the random expanded dataset samples and the target input mixture by adjusting various parameters of the dataset expansion filter based on the values of the frequency band deviation scores.
[0103] 11A and 11B, an exemplary implementation of a machine learning application 1100 will now be described according to one or more implementations. The audio processing systems and methods disclosed herein may be used in conjunction with other audio processing applications that contribute improvements and / or useful functionality to sound post-production editing workflows.
[0104] Referring to block 5.1, a plug-in such as the Avid Audio Extension (AAX) plug-in hosted in a digital audio workstation (DAW) (e.g., a digital audio workstation sold under the trade name PRO TOOLS may be used in one or more implementations) is provided that allows a user to send audio clips from a standalone application, where they may subsequently be processed using a machine learning model. The plug-in can return an arbitrary number of stem splits back to the DAW environment.
[0105] Referring to block 5.2, implementations of the present disclosure also load / receive media used in an application (e.g., an application sold under the trade name JAM LAB with JAM CONNECT or a similar application) that is then processed by a user-selected machine learning recipe. In some implementations, multiple client machines 1 102A-C and processing nodes 1 106A-D with access to client software (e.g., JAM LAB with JAM CONNECT) are configured to access a task manager / database 1104 for access to both the client software and machine learning (ML) application 1100 as disclosed herein.
[0106] Referring to FIG. 12 , an exemplary process flow will now be described. A system running a digital audio workstation 1202 (e.g., PRO TOOLS with JAM CONNECT) is configured to send audio clips to a client application 1208 (e.g., including JAM LAB). Through the client application, a single model may be selected from a list of categories 1210 to process / separate audio sources from the audio clip. A stem type is also selected in step 1212 to form a multi-model recipe. The audio clip and recipe are sent (in step 1214) to a task manager / database 1216, which manages and distributes recipe processing across the user's available processing nodes in step 1218. The client application receives and returns the processed audio clip labeled with the stem and / or model name in step 1206. In some implementations, the client application 1208 may also facilitate selecting stems to form a multi-model recipe 1212 as disclosed herein.
[0107] An implementation of a machine learning multi-model recipe processing system will now be described with reference to Figures 13A-E. In some implementations, a user may desire to separate targeted media into a set of source classes / stems in one step using a selection of source separation models in a particular order and hierarchical combination. A sequential / divergent source separation recipe schema is implemented to process targeted media using one or more source separation models in a sequential / divergent structured order to separate the targeted media into a user-selected set of source classes or stems.
[0108] In the implementation 1300 of FIG. 13A , each step in the recipe represents a source separation model or combination of models targeting a particular source class. The recipe includes appropriately trained models defined according to a recipe schema for performing the steps that result in the user's desired stem output. As illustrated in the exemplary implementation, a user may first choose to separate hiss, followed by voice (continuously separated into vocals and other speech), drums (continuously separated into kick, snare, and other percussion), organ, piano, bass, and other processing. Thus, by processing the targeted media through a pipeline defined by the recipe, the outputs at the various steps may be collected, ultimately separating the targeted media into a set of user-selected source classes or stems.
[0109] Referring to FIG. 13B, an example voice / drums, etc., sequential processing pipeline 1320 is illustrated. In this model, the input mixture is first processed to extract the voice, and the complement includes the drums and other sounds in the mixture. The drum model extracts the drums, and the complement includes the other sounds. In this implementation, the output includes the voice, drums, and other stems. The order in which models are applied when separating source classes using a sequential / divergent separation system may be optimized for higher quality by using an algorithm that assesses the optimal processing order. An example implementation of an optimized processing method 1340 is illustrated in FIG. 13C. For example, an input mixture model with A and B components may be configured to isolate class A, then the remaining B. This may yield different results if the processing order is reversed (e.g., isolating B, then A). The optimized processing method 1340 may operate on an input mixture of A+B by separating A and B in both orders, comparing the results, and selecting the order with the best results (e.g., the results with fewer errors in the separated stems). In various implementations, the optimized processing method 1340 may operate manually, automatically, and / or in a hybrid approach. For example, an estimated optimal order may be pre-established by using a set of ground truth test samples and then creating test mixtures that can be separated using various permutations of stem order; thus, an estimated best-performing model processing order may be established by using an error function that compares the output of these tests against the ground truth test samples.
[0110] An exemplary pipeline 1360 for improving the output fidelity of a sequential / divergent source separation system (e.g., as described previously herein) will now be described with reference to FIG. 13D. For example, output fidelity may be measured using the MOS algorithm, which is capable of measuring the deviation of a stem's output from a particular labeled dataset, such as a collection of speech samples from an individual. In some implementations, such an algorithm may be implemented as a neural network that is pre-trained to either classify sources or measure the deviation of sources from a given dataset.
[0111] A processing recipe (see 5.2.1) may also include a post-processing step after one or more outputs in the pipeline. The post-processing step may include any type of digital signal processing filter / algorithm that cleans up signal artifacts / noise that may have been introduced by previous steps. Pipeline 1360 uses models specifically trained to clean up artifacts (e.g., as described in step 4.2 herein), resulting in significantly improved overall results, especially due to the sequential nature of the recipe processing pipeline.
[0112] 13E, an example implementation 1380 combines models to isolate specific source classes. A sequential / divergent separation processing recipe (see 5.2.1) may include steps in which one or more models are used in combination to extract a source class from a mixture that otherwise could not be fully extracted by only one model trained to target that source class. In the example shown, Drum is separated twice, subsequently summed, and presented as a single "Drum" output stem, accompanied by a complementary "Other" stem.
[0113] An exemplary audio processing system 1400 for implementing the systems and methods disclosed herein will now be described with reference to Figure 14. The audio processing system 1400 includes a logic device 1402, a memory 1404, a communication component 1422, a display 1418, a user interface 1420, and a data storage device 1430.
[0114] Logic device 1402 may include, for example, a microprocessor, a single-core processor, a multi-core processor, a microcontroller, a programmable logic device configured to perform processing operations, a DSP device, one or more memories for storing executable instructions (e.g., software, firmware, or other instructions), a graphics processing unit, and / or any other suitable combination of processing devices and / or memories configured to execute instructions to perform any of the various operations described herein. Logic device 1402 is adapted to interface and communicate with various components of audio processing system 1400, including memory 1404, a communication component 1422, a display 1418, a user interface 1420, and data storage 1430.
[0115] The communications component 1422 may include wired and wireless communications interfaces to facilitate communication with a network or remote system. The wired communications interface may be implemented as one or more physical network or device connection interfaces, such as a cable or other wired communications interface. The wireless communications interface may be implemented as one or more Wi-Fi, Bluetooth, cellular, infrared, radio, and / or other types of network interfaces for wireless communications. The communications component 1422 may include an antenna for wireless communications during operation.
[0116] Display 1418 may include an image display device (e.g., a liquid crystal display (LCD)) or various other types of commonly known video displays or monitors. User interface 1420, in various implementations, may include user input and / or interface devices such as a keyboard, a control panel unit, a graphical user interface, or other user input / output. Display 1418 may operate as both a user input device and a display device, such as, for example, a touchscreen device adapted to receive input signals from a user touching different portions of the display screen.
[0117] Memory 1404 stores program instructions for execution by logic device 1402, including program logic for implementing the systems and methods disclosed herein, including, but not limited to, audio source separation tools 1406, core model operations 1408, machine learning training 1410, trained audio separation model 1412, audio processing application 1414, and self-iteration / training logic 1416. Data used by audio processing system 1400 may be stored in memory 1404 and / or in data storage 1430 and may include machine learning speech dataset 1432, machine learning music / noise dataset 1434, audio stems 1436, audio mixtures 1438, and / or other data.
[0118] In some implementations, one or more processes may be implemented through a remote processing system, such as a cloud platform, which may be implemented as audio processing system 1400 as described herein.
[0119] 1-14 , including various RNNs and models as described herein. Neural network 1500 is implemented as a recurrent neural network, a deep neural network, a convolutional neural network, or other suitable neural network that receives a labeled training data set 1510 to generate audio output 1512 (e.g., one or more audio stems) for each input audio sample. In various implementations, labeled training data set 1510 may include various audio samples and training mixtures as described herein, such as a training data set (FIG. 3 ), an auto-iterative training data set or a culled data set (FIG. 4 ), training mixtures and data sets described according to the training methods described herein (FIGS. 5-13E ), or other training data sets, as appropriate.
[0120] The training process for generating a trained neural network model involves a forward pass through the neural network 1500 to produce an audio stem or other desired audio output 1512. Each data sample is labeled with the desired output of the neural network 1500, which is compared to the audio output 1512. In some implementations, a cost function may be applied to quantify the error in the audio output 1512, and a backward pass through the neural network 1500 may then be used to adjust the neural network coefficients to minimize the output error.
[0121] The trained neural network 1500 may then be tested for accuracy using a subset of the labeled training data 1510 reserved for validation. The trained neural network 1500 may then be implemented as a model in a runtime environment to perform audio source separation as described herein.
[0122] In various implementations, neural network 1500 processes input data (e.g., audio samples) using input layer 1520. In some examples, the input data may correspond to audio samples and / or audio inputs as previously described herein.
[0123] The input layer 1520 includes a plurality of neurons used to condition input audio data for input to the neural network 1500, which may include feature extraction, scaling, sampling rate conversion, and / or the like. The neurons in the input layer 1520 each generate an output that feeds into the input of one or more hidden layers 1530. The hidden layers 1530 include a plurality of neurons that process the output from the input layer 1520. In some embodiments, the neurons in the hidden layer 1530 each generate an output that is then collectively propagated through additional hidden layers that include a plurality of neurons that process the output from the previous hidden layer. The output of the hidden layer 1530 is fed to the output layer 1540. The output layer 1540 includes one or more neurons used to condition the output from the output layer 1540 to generate a desired output. It should be understood that the architecture of neural network 1500 is merely representative and that other architectures are possible, including neural networks with only one hidden layer, neural networks with no input and / or output layers, neural networks with recurrent layers, and / or the like.
[0124] In some embodiments, the input layer 1520, the hidden layer 1530, and / or the output layer 1540 each include one or more neurons. In some embodiments, the input layer 1520, the hidden layer 1530, and / or the output layer 1540 may each include the same or different numbers of neurons. In some embodiments, each neuron takes a combination (e.g., a weighted sum using a trainable weighting matrix W) of its inputs x, adds an optional trainable bias b, and applies an activation function f to generate an output α as shown in the equation α = f(Wx + b). In some embodiments, the activation function f may be a linear activation function, an activation function with upper and / or lower bounds, a log-sigmoid function, a hyperbolic tangent function, a rectified linear unit function, and / or the like. In some embodiments, the neurons may each have the same or different activation functions.
[0125] In some embodiments, neural network 1500 may be trained using supervised learning, where training data combinations include combinations of input data and ground truth (e.g., expected) output data. Differences between the generated audio output 1512 and the ground truth output data (e.g., landmarks) are fed back into neural network 1500 to correct for various trainable weights and biases. In some embodiments, the differences may be fed back using backpropagation techniques using a stochastic gradient descent algorithm and / or the like. In some embodiments, a large set of training data combinations may be presented to neural network 1500 multiple times until an overall cost function (e.g., mean square error based on the difference between each training combination) converges to an acceptable level.
[0126] An example implementation is described below.
[0127] 1. An audio processing system comprising a deep neural network (DNN) trained to separate one or more audio source signals from a single-track audio mixture.
[0128] 2. The audio processing system of Example 1, wherein the DNN is configured to receive a signal input and generate a signal output without time-domain encoding and / or time-domain decoding.
[0129] 3. The audio processing system of Examples 1-2, wherein the DNN is configured to apply a window function.
[0130] 4. The audio processing system of Examples 1-3, wherein the DNN performs an overlap-add process to smooth banding artifacts.
[0131] 5. The audio processing system of Examples 1-4, wherein audio source separation is performed without applying a mask.
[0132] 6. The audio processing system described in Examples 1-5, where the DNN model is trained using a 48 kHz sample rate.
[0133] 7. The audio processing system of Examples 1-6, wherein the signal processing pipeline operates at 48 kHz.
[0134] 8. The audio processing system of any of Examples 1-7, further comprising a separation strength parameter that controls the strength of the separation process applied to the input audio signal.
[0135] 9. The audio processing system of any of Examples 1-8, further comprising a speech training dataset comprising a plurality of labeled speech samples.
[0136] 10. The audio processing system of Examples 1-9, further comprising a non-speech training data set comprising a plurality of labeled music and / or noise data samples.
[0137] 11. An audio processing system as described in Examples 1-10, further comprising a dataset generation module configured to generate labeled audio samples for use in training the DNN model.
[0138] 12. The audio processing system of any one of Examples 1-11, wherein the data set generation module is a self-iterating data set generator.
[0139] 13. The audio processing system of Examples 1-12, wherein the dataset generation module is configured to generate labeled audio samples from the input audio mixture and / or audio source stems output from the DNN.
[0140] 14. The audio processing system of any one of Examples 1-13, further comprising a data loader configured to apply pre / post blend enhancement.
[0141] 15. The audio processing system of Examples 1-14, wherein the DNN is trained at higher than audible frequencies to recognize distinct stems of audio in the lower audible frequency range.
[0142] 16. The audio processing system of any of Examples 1-15, wherein the data loader is configured to apply enhancements such as reverb, filters, and probability parameters.
[0143] 17. The audio processing system of Examples 1-16, further configured to match the sound quality of the target unidentified mixture during training based on the relative signal level.
[0144] 18. An exemplary method comprising:
[0145] processing audio input data using a trained inference model trained for source separation to generate source-separated stems;
[0146] generating a speech dataset from the source-separated stems;
[0147] generating a noisy data set from a source separation system;
[0148] training the inference model using the speech dataset and the noise dataset to generate an updated inference model; A method comprising:
[0149] 19. The method of example 18, further comprising processing audio input data using the updated inference model.
[0150] 20. The method of Examples 18-19, further comprising the step of iteratively updating the updated inference model.
[0151] 21. The method of Examples 19-20, wherein the training dataset is curated to include samples that approximate the audio source.
[0152] 22. The method of Examples 19-21, further comprising a hierarchical mix bus schema.
[0153] 23. The method of Examples 19-22, in which the inference model is trained using a multi-audio source mixture.
[0154] 24. The method of Examples 19-23, wherein the inference model is trained using foreground speech, background speech, and / or distant speech.
[0155] 25. The method of Examples 19-24, wherein the inference model is trained at a first sample rate and upscaled to a higher sample rate.
[0156] 26. The method of Examples 19-25, further comprising the step of post-processing the separated audio source stems to remove artifacts introduced by the source separation process.
[0157] 27. The method of any one of Examples 19-26, wherein the source separation system includes a separated source signal and a remaining complementary signal.
[0158] 28. The method of Examples 19-27, wherein the artifacts introduced during the source separation process include clicks, harmonic distortion, ghosting, and / or broadband noise.
[0159] 29. The method of Examples 19-28, wherein the fine-tuning process includes a user-guided self-iterative process.
[0160] 30. The method of Examples 19-29, further comprising facilitating user-guided extensions for fine-tuning of training.
[0161] 31. A system comprising: an audio input configured to receive an audio input stream comprising a mixture of audio signals generated from a plurality of audio sources; a trained audio source separation model configured to receive the audio input stream and generate a plurality of generated audio stems, the generated plurality of audio stems corresponding to one or more of the plurality of audio sources; and a self-iterative training system configured to update the trained audio source separation model to an updated audio source separation model based at least in part on the generated plurality of audio stems, the updated audio source separation model configured to reprocess the audio input stream and generate a plurality of improved audio stems.
[0162] 32. The system described in Example 31, wherein the audio input stream comprises one or more single-track audio mixtures, and the trained audio source separation model comprises a neural network trained to separate one or more audio source signals from the one or more single-track audio mixtures.
[0163] 33. The system of Examples 31-32, wherein the neural network is configured to perform audio source separation without applying a mask.
[0164] 34. The system of Examples 31-33, further comprising a training dataset comprising labeled source audio data and labeled noise audio data, wherein the trained audio source separation model is trained using the training dataset to generate a general source separation model.
[0165] 35. The system of Examples 31-34, wherein at least a subset of the generated plurality of audio stems is culled based on a threshold metric and added to a training dataset to form a culled dynamically evolving dataset, and the culled dynamically evolving dataset is used to train an updated audio source separation model.
[0166] 36. The system described in Examples 31-35, wherein the self-iterative training system is further configured to calculate a first quality metric associated with the generated plurality of audio stems, the first quality metric providing a first performance measure of the trained audio source separation model, and the self-iterative training system is further configured to calculate a second quality metric associated with the improved audio stems, the second quality metric providing a second performance measure of the updated audio source separation model, the second quality metric exceeding the first quality metric.
[0167] 37. The system described in Examples 31-36, wherein the trained audio source separation model is trained using a training dataset comprising a plurality of datasets, each of the plurality of datasets comprising labeled audio samples configured to train the system to address a source separation problem.
[0168] 38. The system of Examples 31-37, wherein the plurality of datasets comprises a speech training dataset comprising a plurality of labeled speech samples and / or a non-speech training dataset comprising a plurality of labeled music and / or noise data samples.
[0169] 39. The system of Examples 31-38, wherein the self-repetitive training system further comprises a self-repetitive data set generation module configured to generate labeled audio samples from the generated plurality of audio stems.
[0170] 40. The system of Examples 31-39, wherein the multiple enhanced audio stems are generated using a hierarchical branching sequence, including a step of separating the source signal and the remaining complementary signal.
[0171] 41. A method comprising: receiving an audio input stream comprising a mixture of audio signals generated from a plurality of audio sources; using a trained audio source separation model configured to receive the audio input stream to generate a plurality of generated audio stems corresponding to one or more of the plurality of audio sources; updating the trained audio source separation model to an updated audio source separation model based, at least in part, on the generated plurality of audio stems using a self-iterative training process; and reprocessing the audio input stream using the updated audio source separation model to generate a plurality of improved audio stems.
[0172] 42. The method described in Example 41, wherein the audio input stream comprises one or more single-track audio mixtures, and the trained audio source separation model comprises a neural network that is trained to separate one or more audio source signals from the one or more single-track audio mixtures.
[0173] 43. The method of any one of Examples 41-42, wherein the neural network is configured to perform audio source separation without applying a mask.
[0174] 44. The method of any one of Examples 41-43, further comprising providing a training dataset comprising labeled source audio data and labeled noise audio data, and using the training dataset to train an audio source separation model trained to generate a general source separation model.
[0175] 45. The method of any one of Examples 41-44, further comprising the steps of: adding at least a subset of the generated plurality of audio stems to a training dataset to generate a dynamically evolving dataset; culling the dynamically evolving dataset based on a threshold metric; and training an updated audio source separation model using the culled dynamically evolving dataset.
[0176] 46. The method described in Examples 41-45, wherein the self-iterative training process further includes the steps of calculating a first quality metric associated with the generated plurality of audio stems, the first quality metric providing a first performance measure of the trained audio source separation model; calculating a second quality metric associated with the improved audio stems, the second quality metric providing a performance measure of the updated audio source separation model; and comparing the second quality metric with the first quality metric and confirming that the second quality metric exceeds the first quality metric.
[0177] 47. The method described in Examples 41-46, wherein the trained audio source separation model is trained using a training dataset comprising a plurality of datasets, each of the plurality of datasets comprising labeled audio samples configured to train the audio source separation model to address a different source separation problem.
[0178] 48. The method of any one of Examples 41-47, wherein the plurality of datasets comprises a speech training dataset comprising a plurality of labeled speech samples and / or a non-speech training dataset comprising a plurality of labeled music and / or noise data samples.
[0179] 49. The method of any one of Examples 41-48, wherein the self-repetitive training process further includes generating labeled audio samples from a plurality of audio stems generated for the self-repetitive data set.
[0180] 50. The method of any one of Examples 41-49, further comprising generating a plurality of enhanced audio stems using a hierarchical divergent sequence, including separating the source signal and the remaining complement signal.
[0181] Where applicable, the various implementations provided by the present disclosure can be implemented using hardware, software, or a combination of hardware and software. Also, where applicable, the various hardware and / or software components described herein can be combined into composite components comprising software, hardware, and / or both without departing from the spirit of the present disclosure. Where applicable, the various hardware and / or software components described herein can be separated into subcomponents comprising software, hardware, or both without departing from the spirit of the present disclosure.
[0182] Software according to the present disclosure, such as non-transitory instructions, program code, and / or data, can be stored on one or more non-transitory machine-readable media. It is also contemplated that the software identified herein can be implemented using one or more general-purpose or special-purpose computers and / or computer systems, networked and / or otherwise. Where applicable, the ordering of various steps described herein can be changed, combined into composite steps, and / or separated into substeps to provide the features described herein. The above-described implementations illustrate the invention, but do not limit it. It should also be understood that numerous modifications and variations are possible in accordance with the principles of the invention. The scope of the invention, therefore, is defined only by the following claims.
Claims
1. A system, comprising: a memory component that stores machine-readable instructions; Logical Devices and Equipped with The logical device is a trained audio source separation model configured to receive audio input samples comprising a single-track mixture of audio signals generated from a plurality of audio sources and to generate a plurality of audio stems, the plurality of audio stems corresponding to one or more audio sources of the plurality of audio sources; a self-iterative training system configured to perform multiple training iterations, the training iteration including generating a new audio source separation model based at least in part on a training dataset comprising a subset of the generated plurality of audio stems from a previous training iteration, the subset of the generated plurality of audio stems comprising one or more of a portion of a stem, portions of a stem, a complete stem, multiple stems, or a combination of a complete stem and one or more portions of another stem, the new audio source separation model generated in each iteration being increasingly specific to the mixture of audio signals in the audio input sample; configured to execute the machine-readable instructions on The new audio source separation model is configured to generate a plurality of improved audio stems by reprocessing the audio input stream.
2. The system described in claim 1, wherein the trained audio source separation model comprises a neural network trained to separate one or more audio source signals from the single-track mixture of audio signals.
3. The system described in claim 2, wherein the neural network is configured to perform audio source separation without applying a mask.
4. The system described in claim 1, further comprising a training dataset including labeled source audio data and labeled noise audio data, and the trained audio source separation model is initially trained using the training dataset to generate a general source separation model.
5. The system described in claim 4, wherein during each training iteration, at least a subset of the plurality of audio stems is culled based on a threshold metric and added to the training dataset from the previous training iteration to form a culled dynamically evolving dataset, and the culled dynamically evolving dataset is used to train the new audio source separation model.
6. The system of claim 1, wherein the trained audio source separation model is trained using a training dataset including a plurality of datasets, each of the plurality of datasets comprising labeled audio samples configured to train the system to address source separation associated with identified sources.
7. The system described in claim 6, wherein the multiple datasets include a speech training dataset comprising a plurality of labeled speech samples, and / or a non-speech training dataset comprising a plurality of labeled music and / or noise data samples.
8. The system described in claim 1, wherein the self-repetitive training system further comprises a self-repetitive dataset generation module configured to generate labeled audio samples from the generated plurality of audio stems.
9. The system described in claim 1, wherein the multiple enhanced audio stems are generated using a hierarchical branching sequence that includes separating a source signal from a remaining complementary signal.
10. A method, comprising: receiving an audio input stream comprising a mixture of audio signals generated from a plurality of audio sources; generating a plurality of generated audio stems corresponding to one or more audio sources of the plurality of audio sources using a trained audio source separation model configured to receive the audio input stream; updating the trained audio source through multiple training iterations, wherein a training iteration includes generating a new audio source separation model based at least in part on a subset of the generated multiple audio stems derived from a previous training iteration, wherein the subset of the generated multiple audio stems comprises one or more of a portion of a stem, multiple portions of a stem, a complete stem, multiple stems, or a combination of a complete stem and one or more portions of another stem; generating a plurality of improved audio stems by reprocessing the audio input stream using the new audio source separation model; and A method comprising:
11. The method described in claim 10, wherein the audio input stream comprises one or more single-track audio mixtures, and the trained audio source separation model comprises a neural network trained to separate one or more audio source signals from the one or more single-track audio mixtures.
12. The method of claim 11, wherein the neural network is configured to perform audio source separation without applying a mask.
13. The method comprising: providing a training data set including labeled source audio data and labeled noisy audio data; generating a general source separation model by training the trained audio source separation model using the training dataset; The method of claim 10 further comprising:
14. The method comprising: generating a dynamically evolving dataset by adding at least a subset of the generated plurality of audio stems to the training dataset; culling the dynamically evolving dataset based on a threshold metric; and training the new audio source separation model using the culled dynamically evolving dataset; and 14. The method of claim 13, further comprising:
15. The method described in claim 10, wherein the trained audio source separation model is trained using a training dataset including a plurality of datasets, each of the plurality of datasets comprising labeled audio samples configured to train the audio source separation model to address a different source separation problem.
16. The method of claim 15, wherein the plurality of datasets comprises a speech training dataset comprising a plurality of labeled speech samples, and / or a non-speech training dataset comprising a plurality of labeled music and / or noise data samples.
17. The method described in claim 16, wherein updating the trained audio source separation model further comprises generating labeled audio samples from the generated plurality of audio stems for a self-repeating dataset.
18. The method of claim 10, further comprising generating the multiple improved audio stems using a hierarchical branching sequence that includes separating the source signal and a remaining complementary signal.