Audio source separation processing pipeline system and method

The audio source separation system uses machine learning models to refine low-quality single-track recordings, addressing the challenge of generating high-fidelity audio stems from noisy, older recordings by isolating and enhancing audio components.

JP2026091907APending Publication Date: 2026-06-04WINGNUT FILMS PROD LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
WINGNUT FILMS PROD LTD
Filing Date
2026-03-19
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing audio source separation techniques are not optimized to generate high-quality audio stems from low-quality, noisy single-track recordings, particularly in older sound recordings.

Method used

An audio source separation system using machine learning models to separate single-track recordings into high-fidelity stems, with a sequential audio source separation model and neural networks to refine and remove artifacts.

Benefits of technology

The system effectively isolates and enhances audio components from low-quality recordings, reducing artifacts and noise to produce high-fidelity audio stems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026091907000002
    Figure 2026091907000002
  • Figure 2026091907000003
    Figure 2026091907000003
  • Figure 2026091907000004
    Figure 2026091907000004
Patent Text Reader

Abstract

To provide a system and method for audio source isolation. [Solution] A system and method for audio source separation includes the steps of receiving a single-track audio input sample having an unknown mixture of audio signals generated from a plurality of audio sources, and separating one or more of the audio sources from the single-track audio input sample using a sequential audio source separation model. The step of separating one or more of the audio sources may include defining a processing recipe that includes a plurality of source separation processes configured to receive the audio input mixture and output one or more separated source signals and the remaining complementary signal mixture, and processing the single-track audio input sample according to the processing recipe to generate a plurality of audio stems separated from the unknown mixture of audio signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - reference to Related Applications) This disclosure claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 272,650, filed Oct. 27, 2021, and claims the priority of U.S. Patent Application No. 17 / 848,341, filed Jun. 23, 2022, both of which are hereby incorporated by reference in their entireties.

[0002] This disclosure generally relates to systems and methods for audio source separation, and more particularly, to systems and methods for separating and enhancing audio source signals from audio mixtures such as single - track audio mixtures.

Background Art

[0003] Audio mixing is the process of combining multiple audio recordings to produce an optimized mixture for playback in one or more desired audio formats such as monoral, stereo, or surround sound. In applications that require high - quality audio production such as audio generation for music and movies, audio mixtures are generally produced by mixing separate high - quality recordings. These separate recordings often occur in a controlled environment such as a recording studio with optimized acoustic effects and high - quality recording equipment.

[0004] Often, a portion of the source audio is of low quality and / or may include a mixture of desired audio sources and unwanted noise. In modern audio post - production, it is common to rerecord audio when the original recording lacks the desired quality. For example, in music recording, vocal or instrument tracks can be recorded and mixed with previous recordings. In audio post - production for movies, it is common to invite actors to a studio, rerecord their dialogue, and add other audio (e.g., sound effects, music) to the mix.

[0005] However, in some applications, it is desirable to faithfully convert the original audio source into a high-quality audio mix. For example, film, music, television broadcasts, and other audio recordings can date back over 100 years. The source audio may have been recorded with older, lower-quality equipment and may contain a low-quality mixture of the desired audio and noise. In many recordings, a single-track / mono audio mix is ​​the only audio source available to produce an optimized mixture for playback on a modern sound system.

[0006] One approach to processing audio mixtures is to separate them into distinct sets of audio source components and generate separate audio stems for each component of the audio mixture. For example, a music recording can be separated into vocal, guitar, bass, and drum components. Each of these distinct components can then be enhanced and blended to optimize playback.

[0007] However, existing audio source separation techniques are not optimized to generate the high-quality audio stems required to produce high-fidelity output for the music and film industries. Audio source separation is particularly difficult when the audio source is of lower quality from older sound recordings and contains a noisy mixture of single tracks.

[0008] Considering the foregoing, there is a continuing need for improved audio source isolation systems and methods, particularly for generating high-fidelity audio from low-quality audio sources.

[0009] Addressing at least some of the aforementioned disadvantages is the objective of at least the preferred embodiment. An additional or alternative objective is to provide the public with a useful alternative to the conventional technique. [Overview of the Initiative] [Means for solving the problem]

[0010] Improved audio source separation systems and methods are disclosed herein. Various implementations provide an audio source separation system configured to separate a single-track audio recording into various audio components, such as speech and individual instruments, into high-fidelity stems (e.g., discrete or grouped collections of mixed audio sources).

[0011] In some implementations, the audio source separation system includes a first machine learning model trained to separate a single-track audio recording into stems containing utterances, complements, and artifacts such as "clicks." Additional machine learning models may then be used to refine the utterance stems by removing processing artifacts from the utterances and / or fine-tuning the first machine learning model.

[0012] As used herein, the term "comprising" means "consisting at least in part of." When interpreting each word herein that contains the term "comprising," other characteristics may also exist besides those preceded or grouped by that term. Related terms such as "comprise" and "comprises" are also interpreted in the same manner.

[0013] In various implementations, the method includes the steps of: receiving a single-track audio input sample comprising an unknown mixture of audio signals generated from multiple audio sources; defining a processing recipe comprising multiple source separation processes configured to receive the audio input mixture and output one or more isolated source signals and the remaining complementary signal mixture; and separating one or more audio sources from a single-track audio input sample using a sequential audio source separation model, which includes the steps of: processing the single-track audio input sample according to the processing recipe to generate multiple audio stems separated from the unknown mixture of audio signals.

[0014] The method may further define a processing recipe by: processing a single-track audio input sample using a first processing order of multiple source separation processes for separating audio stems to generate a first set of audio stems; processing a single-track audio input sample using a second processing order of multiple source separation processes different from the first processing order for separating audio stems to generate a second set of audio stems; and evaluating the first set of audio stems and the second set of audio stems to determine which of the first and second processing orders should be included in the processing recipe.

[0015] The processing recipe of this method may further include a step of post-processing multiple audio stems to remove and / or reduce artifacts introduced by a sequential audio source separation model, the post-processing step of processing the multiple audio stems through one or more neural networks trained to remove and / or reduce artifacts, including clicks, harmonic distortion, ghosting, and / or broadband noise.

[0016] The processing recipe of this method may further include a step of performing multiple source separation processes in a sequential branching processing order, where a first source separation process is configured to receive a single-track audio input sample and output one or more source separation signals and the remaining complementary signal mixture, and each subsequent separation process is configured to receive the output signal from the preceding source separation process, with each node of the processing recipe configured to separate a particular source class from the input audio mixture.

[0017] The processing recipe of this method may further include the steps of combining two or more of a plurality of audio stems having a common source class generated by different source separation processes in the processing recipe, and / or post-processing the plurality of audio stems by processing one or more of the plurality of audio stems and applying a neural network model trained to eliminate artifacts on the audio stems separated according to the processing recipe, thereby reducing artifacts and / or noise introduced by the sequential audio source separation model.

[0018] At least one source separation process of this method may further include a step of separating and outputting a mixture of a speech source and the remaining complement comprising music and noise.

[0019] The method may further include a step of evaluating output audio stems from multiple sequential branching processing sequences and assessing the optimized processing sequence, the evaluation step of processing the input mixture to separate a first source class and a first remaining complement, then processing the first remaining complement, separating a second source class, and generating a first set of audio output stems, the evaluation step of further processing the input mixture to separate a second source class and a second remaining complement, then processing the second remaining complement, separating the first source class, and generating a second set of audio output stems. The method may further include a step of comparing the first set of audio output stems and the second set of audio output stems to determine the processing sequence that generates higher quality audio output stems.

[0020] The step of separating one or more audio sources from a single-track audio input sample using the sequential audio source separation model of this method may further include a user-inducing and / or self-repeating process configured to progressively separate sources from an unknown mixture of audio signals by selecting audio enhancement and / or audio enhancement parameters, fine-tuning the sequential audio source separation model to enable matching source signals within an unknown mixture of audio signals, and identifying source classes within the unknown mixture of audio signals and identifying a general separation model for use in a processing recipe.

[0021] The processing recipe of the recipe may further comprise a set of source classes and output stems and corresponding source separation models, arranged in a hierarchical branching sequence, wherein at least one class of sources is separated into a plurality of source class separation stems, and / or the processing recipe includes a step of generating a mixture of the plurality of source class separation stems as a source class output.

[0022] The method may further include the step of training a sequential audio source separation model using a single-track audio input sample by providing a single-track audio input sample to a general source separation model, generating a plurality of initial audio stems corresponding to one or more of a plurality of audio sources, and retraining the general source separation model using at least partially one or more of the plurality of initial audio stems, wherein the general source separation model comprises a plurality of neural network models, each neural network model configured to receive a single-channel audio input sample and output one or more source-separated audio stems comprising a source class and a mixture of the remaining complementary signals.

[0023] The step of training a sequential audio source separation model of this method may further include a step of training multiple neural network models in an order defined by a processing recipe, and / or the processing recipe comprises a hierarchical branching sequence in which each branch comprises one or more of the multiple neural network models. The step of training a sequential audio source separation model of this method may further include, at least in part, a step of re-training the sequential audio source separation model based on multiple audio stems generated from a prior iteration of the sequential audio source separation model, a step of determining a metric associated with the source separation audio stems and evaluating one or more of the multiple audio stems by comparing the metric with one or more threshold parameters, and / or a step of adding one or more source separation audio stems to a training dataset to train the sequential audio source separation model based on the evaluation step.

[0024] The training step may further include generating a training dataset of artifacts generated by a processing recipe to train one or more neural networks, receiving a source separation output with artifacts, and generating an improved output in which one or more artifacts are reduced and / or removed according to the processing recipe.

[0025] In various implementations, the system defines a processing recipe including a memory component storing machine-readable instructions and a plurality of source separation processes configured to receive an audio input mixture and output one or more separated source signals and a remaining complementary signal mixture, and processes a single-track audio input sample according to the processing recipe to generate a plurality of audio systems separated from an unknown mixture of audio signals. A logic device configured to execute machine-executable instructions to separate one or more audio sources from a single-track audio input sample comprising an unknown mixture of audio signals generated from a plurality of audio sources.

[0026] The logic device may further be configured to define a processing recipe by processing a single-track audio input sample using a first processing order of a plurality of source separation processes for separating audio systems to generate a first set of audio systems, and using a second processing order of a plurality of source separation processes different from the first processing order for separating audio systems to process the single-track audio input sample to generate a second set of audio systems, and evaluating the first set of audio systems and the second set of audio systems to determine which of the first processing order and the second processing order should be included in the processing recipe.

[0027] The logic device may further be configured to define a processing recipe by post - processing a plurality of audio systems and removing and / or reducing artifacts introduced by a sequential audio source separation model, the post - processing step including processing the plurality of audio systems through one or more neural networks trained to remove and / or reduce artifacts including click sounds, harmonic distortion, ghosting, and / or broadband noise.

[0028] The system may further be defined in that at least one source separation process includes separating and outputting a mixture of a speech source and the remaining complement including music and noise.

[0029] The logic device of the system may further be configured to define processing by performing a plurality of source separation processes in a sequential branched processing order, where the first source separation process is configured to receive a single - track audio input sample and output one or more source separation signals and a mixture of remaining complement signals, and each subsequent separation process is configured to receive the output signal from a previous source separation process, and each node of the processing recipe is configured to separate a specific source class from the input audio mixture.

[0030] The logic device may further be configured to evaluate output audio stems from multiple sequential branching processing sequences and assess the optimized processing sequence by: processing an input mixture to separate a first source class and a first remaining complement, then processing the first remaining complement, separating a second source class, and generating a first set of audio output stems; processing an input mixture to separate a second source class and a second remaining complement, then processing the second remaining complement, separating the first source class, and generating a second set of audio output stems; and comparing the first set of audio output stems with the second set of audio output stems to determine a processing sequence that generates higher quality audio output stems.

[0031] This summary is provided to introduce in a simplified form a set of concepts that will be further described below in the detailed description. This summary is not intended to identify any important or essential features of the claimed subject matter, nor to limit the scope of the claimed subject matter. A broader presentation of the features, details, usefulness, and advantages of the method as defined in the claims is provided in the following description of various implementations of the disclosure and illustrated in the accompanying drawings.

[0032] Where patent specifications, other external documents, or other sources of information are referenced herein, this is generally intended to provide a context for discussing the features of the invention. Unless otherwise specifically stated, references to such external documents or sources of information shall not be construed as an admission that such documents or sources of information are prior art or form part of the common general knowledge in the art in any jurisdiction. The present invention provides, for example, the following items: (Item 1) It is a method, Receiving a single-track audio input sample containing an unknown mixture of audio signals generated from multiple audio sources, Defining a processing recipe that includes multiple source separation processes configured to receive an audio input mixture and output one or more separated source signals and the remaining complementary signal mixture, The process involves processing the single-track audio input sample according to the processing recipe described above, generating multiple audio stems separated from the unknown mixture of the audio signals, and Using a sequential audio source separation model that includes the following, separate one or more of the audio sources from the single-track audio input sample: Methods that include... (Item 2) Defining the aforementioned processing recipe means that Using a first processing sequence of the multiple source separation processes for separating audio stems, the single-track audio input sample is processed to generate a first set of audio stems, Processing the single-track audio input sample using a second processing order of the multiple source separation processes, which differs from the first processing order for separating audio stems, to generate a second set of audio stems, Evaluate the first set of audio stems and the second set of audio stems, and determine which of the first and second processing sequences should be included in the processing recipe. The method described in item 1, including the method described in item 1. (Item 3) The method according to item 1 or item 2, wherein the processing recipe further comprises post-processing the plurality of audio stems to remove and / or reduce artifacts introduced by the sequential audio source separation model, wherein the post-processing comprises processing the plurality of audio stems through one or more neural networks trained to remove and / or reduce artifacts, including clicks, harmonic distortion, ghosting, and / or broadband noise. (Item 4) The method according to any one of items 1-3, wherein at least one source separation process includes separating and outputting a mixture of a speech source and the remaining complement comprising music and noise. (Item 5) The processing recipe further includes the step of executing the plurality of source separation processes in a sequential branching processing sequence, each subsequent separation process configured to receive the output signal from the preceding source separation process, wherein a first source separation process is configured to receive the single-track audio input sample and output one or more source separation signals and a mixture of the remaining complementary signals, and each subsequent separation process is configured to receive the output signal from the preceding source separation process. Each node in the aforementioned processing recipe is configured to isolate a specific source class from the input audio mixture. The method described in item 1. (Item 6) This involves evaluating output audio stems from multiple sequential branching processing sequences and assessing the optimized processing sequence. The evaluation involves processing the input mixture, separating the first source class and the first remaining complement, then processing the first remaining complement, separating the second source class, and generating a first set of audio output stems. The evaluation further includes processing the input mixture to separate the second source class and the second remaining complement, then processing the second remaining complement to separate the first source class and generate a second set of audio output stems. The first set of audio output stems and the second set of audio output stems are compared to determine a processing order that generates higher quality audio output stems. The method described in item 5, further including the method described in item 5. (Item 7) The aforementioned processing recipe is: Combining two or more of the plurality of audio stems having a common source class generated by different source separation processes within the processing recipe, and / or, To reduce artifacts and / or noise introduced by the sequential audio source separation model by processing one or more of the aforementioned audio stems and applying a neural network model trained to eliminate artifacts on the separated audio stems according to the processing recipe. The method according to any one of items 1-6, comprising post-processing the plurality of audio stems including the above. (Item 8) Using a sequential audio source separation model, separating one or more of the audio sources from the single-track audio input sample further involves: Select audio extensions and / or audio extension parameters to fine-tune the sequential audio source separation model and enable matching source signals within an unknown mixture of audio signals, To identify the source class within the unknown mixture of the audio signals and to identify a general separation model for use in the processing recipe. The method according to any one of items 1-7, comprising a user-inducing and / or self-repeating process configured to progressively isolate a source from an unknown mixture of audio signals containing the audio signal. (Item 9) The processing recipe comprises a set of source classes and output stems and corresponding source separation models arranged in a hierarchical branching sequence, At least one class of the source is separated into a plurality of source class separation stems, and the processing recipe includes generating a mixture of the plurality of source class separation stems as the source class output. The method described in item 8. (Item 10) The method further includes providing the single-track audio input sample to a general source separation model, generating a plurality of initial audio stems corresponding to one or more of the plurality of audio sources, and training the sequential audio source separation model using the single-track audio input sample by retraining the general source separation model using at least partially one or more of the plurality of initial audio stems. The aforementioned general source separation model comprises multiple neural network models, each neural network model configured to receive a single-channel audio input sample and output one or more source-separated audio stems comprising a source class and a mixture of the remaining complementary signals. The method described in any one of items 1-9. (Item 11) Training the sequential audio source separation model further includes training the plurality of neural network models in the order defined by the processing recipe, The processing recipe comprises a hierarchical branching sequence in which each branch comprises one or more of the plurality of neural network models. The method described in item 10. (Item 12) The method of item 11, wherein training the sequential audio source separation model further comprises, at least in part, re-training the sequential audio source separation model based on the plurality of audio stems generated from previous iterations of the sequential audio source separation model. (Item 13) Training the aforementioned sequential audio source separation model further involves, Determining a metric associated with the source-separated audio stem, and evaluating one or more of the multiple audio stems by comparing the metric with one or more threshold parameters, Based on the evaluation described above, one or more source-separated audio stems are added to the training dataset in order to train the sequential audio source separation model. The method described in item 12, including the method described in item 12. (Item 14) The method according to any one of items 10-13, further comprising training the sequential audio source separation model to generate a training dataset of artifacts generated by the processing recipe to train one or more neural networks, receiving source separation outputs with artifacts, and generating improved outputs in which one or more artifacts are mitigated and / or removed according to the processing recipe. (Item 15) It is a system, A memory component that stores machine-readable instructions, A logical device, wherein the logical device executes the machine-readable instruction, Defining a processing recipe that includes multiple source separation processes configured to receive an audio input mixture and output one or more separated source signals and the remaining complementary signal mixture, The process involves processing a single-track audio input sample according to the aforementioned processing recipe to generate multiple audio stems separated from an unknown mixture of audio signals. This involves using a sequential audio source separation model to separate one or more audio sources from the single-track audio input sample, which comprises an unknown mixture of audio signals generated from multiple audio sources. Logical devices and A system equipped with these features. (Item 16) The aforementioned logical device further, Using a first processing sequence of the multiple source separation processes for separating audio stems, the single-track audio input sample is processed to generate a first set of audio stems, Processing the single-track audio input sample using a second processing order of the multiple source separation processes, which differs from the first processing order for separating audio stems, to generate a second set of audio stems, Evaluate the first set of audio stems and the second set of audio stems, and determine which of the first and second processing sequences should be included in the processing recipe. The system described in item 15, configured to define the processing recipe accordingly. (Item 17) The aforementioned logical device further, The system is configured to define a processing recipe by post-processing the aforementioned multiple audio stems and removing and / or mitigating artifacts introduced by the sequential audio source separation model, The post-processing described above includes processing the plurality of audio stems through one or more neural networks trained to remove and / or reduce artifacts including clicks, harmonic distortion, ghosting, and / or broadband noise. The system described in item 15 or item 16. (Item 18) A system according to any one of items 15–17, comprising at least one source separation process for separating and outputting a mixture of a speech source and the remaining complement comprising music and noise. (Item 19) The aforementioned logical device further, A first source separation process is configured to receive the single-track audio input sample and output one or more source separation signals and a mixture of the remaining complementary signals, and each subsequent separation process is configured to execute the multiple source separation processes in a sequential branching processing order in which each subsequent separation process receives an output signal from a preceding source separation process. The process is defined by the following: Each node in the aforementioned processing recipe is configured to isolate a specific source class from the input audio mixture. A system as described in any one of items 15-18. (Item 20) The aforementioned logical device further, The input mixture is processed to separate the first source class and the first remaining complement, then the remaining complement is processed to separate the second source class and generate a first set of audio output stems. Processing the input mixture to separate the second source class and the second remaining complement, then processing the second remaining complement to separate the first source class and generate a second set of audio output stems, The first set of audio output stems and the second set of audio output stems are compared to determine a processing order that generates higher quality audio output stems. The system described in item 19 is configured to evaluate output audio stems from multiple sequential branching processing sequences and to assess the optimized processing sequence. [Brief explanation of the drawing]

[0033] Aspects of this disclosure and their benefits can be better understood by referring to the following drawings and the subsequent detailed description. Similar reference numbers are used to identify similar elements illustrated in one or more of the drawings, and it should be understood that the illustrations therein are intended to illustrate, and not limit, an implementation of this disclosure. Components in the drawings are not necessarily to scale, and instead the emphasis is on clearly illustrating the principles of this disclosure.

[0034] [Figure 1] Figure 1 illustrates an audio source isolation system and process with one or more implementations.

[0035] [Figure 2] Figure 2 illustrates the elements related to the system and process in Figure 1, with one or more implementations.

[0036] [Figure 3] Figure 3 is a schematic diagram illustrating machine learning datasets and training data loaders in one or more implementations.

[0037] [Figure 4] Figure 4 illustrates an exemplary machine learning training system, including a self-repeating dataset generation loop, with one or more implementations.

[0038] [Figure 5] Figure 5 illustrates exemplary operation of a data loader for use in training machine learning systems, with one or more implementations.

[0039] [Figure 6] Figure 6 illustrates exemplary machine learning training methods using one or more implementations.

[0040] [Figure 7] Figure 7 illustrates an exemplary machine learning training method, including a training mixture example with one or more implementations.

[0041] [Figure 8] Figure 8 illustrates exemplary machine learning processing with one or more implementations.

[0042] [Figure 9] Figure 9 illustrates an exemplary post-processing model configured to clean up artifacts introduced by machine learning processing, with one or more implementations.

[0043] [Figure 10] Figure 10 illustrates an exemplary user-inducing self-repeating training loop with one or more implementations.

[0044] [Figure 11A]Figure 11 includes Figures 11A and 11B, illustrating exemplary machine learning applications with one or more implementations. [Figure 11B] Figure 11 includes Figures 11A and 11B, illustrating exemplary machine learning applications with one or more implementations.

[0045] [Figure 12] Figure 12 illustrates an exemplary machine learning processing application with one or more implementations.

[0046] [Figure 13A] Figure 13 includes Figures 13A, 13B, 13C, 13D, and 13E, illustrating one or more embodiments of a multi-model recipe processing system with one or more implementations. [Figure 13B] Figure 13 includes Figures 13A, 13B, 13C, 13D, and 13E, illustrating one or more embodiments of a multi-model recipe processing system with one or more implementations. [Figure 13C] Figure 13 includes Figures 13A, 13B, 13C, 13D, and 13E, illustrating one or more embodiments of a multi-model recipe processing system with one or more implementations. [Figure 13D] Figure 13 includes Figures 13A, 13B, 13C, 13D, and 13E, illustrating one or more embodiments of a multi-model recipe processing system with one or more implementations. [Figure 13E] Figure 13 includes Figures 13A, 13B, 13C, 13D, and 13E, illustrating one or more embodiments of a multi-model recipe processing system with one or more implementations.

[0047] [Figure 14] Figure 14 illustrates an exemplary audio processing system with one or more implementations.

[0048] [Figure 15] Figure 15 illustrates an exemplary neural network that may be used in one or more implementations of the implementations shown in Figures 1-14. [Modes for carrying out the invention]

[0049] Detailed explanation The following description will explain various implementations. For explanatory purposes, specific configurations and details are provided to give a thorough understanding of the implementation. However, it will also be apparent to those skilled in the art that the implementation can be practiced without specific details. Furthermore, well-known features may be omitted or simplified to avoid obscuring the described implementation.

[0050] Improved audio source separation systems and methods are disclosed herein. In various implementations, an audio source separation system is provided which a single-track (e.g., undifferentiated) audio recording is configured to separate and divide various audio components, such as speech and instruments, into high-fidelity stems, i.e., discrete or grouped collections of mixed audio sources. In various implementations, the single-track audio recording contains an unidentified audio mixture (e.g., the audio sources, recording environment, and / or other aspects of the audio mixture are unknown to the audio source separation system), and the audio source separation system and method are adapted to identify and / or separate audio sources from the unidentified audio mixture in a self-repetitive training and fine-tuning process.

[0051] The systems and methods disclosed herein may be implemented on at least one computer-readable medium that carries instructions, which, when executed by at least one processor, cause at least one processor to perform any of the method steps disclosed herein. Some implementations relate to a computer system that includes at least one processor and memory that stores instructions, which, when executed by at least one processor, causes at least one processor to perform any of the method steps disclosed herein. In various implementations, the models described herein may be implemented as stored data and software modules and / or code that acts on the stored data.

[0052] In some cases, the step of training a model involves processing one or more data structures to form a new data structure that can be accessed by components of the computer system and used as a model. For example, an artificial intelligence system may comprise a computer having one or more processors, program code memory, writable data memory, and several inputs / outputs. The writable data memory may hold several data structures corresponding to trained or untrained models. Such data structures may represent one or more layers of nodes in a neural network, links between nodes in different layers, and weights relating to at least some of the links between nodes. In other cases, different types of data structures may represent the model.

[0053] In some cases, when referring to the steps of training a model, feeding a model, and / or having a model take in inputs and provide outputs, this may refer to the actions of a computer capable of reading writable data memory, which contains the model and executes program code to work with the model. For example, a model may be trained on a set of training data, which may be the training examples themselves and / or the training examples and their corresponding ground truth. Once trained, the model may be available to make decisions about the examples provided to the model. This may be done by the computer receiving input data representing the examples, performing a process with the examples and the model, and outputting output data representing and / or indicating decisions made by or based on the model.

[0054] In a very specific embodiment, an artificial intelligence system may have a processor that reads a large number of photographs of cars and reads ground truth data indicating that "these are cars." The processor may also read a large number of photographs of streetlights and similar objects and read ground truth data indicating that "these are not cars." In some cases, the model is trained on the input data itself without being provided with ground truth data. The result of such processing may be a trained model. The artificial intelligence system can then be provided with an image that, in the state in which the model is being trained, does not have any indication of whether it is an image of a car, and can output an indication of the decision of whether it is an image of a car.

[0055] To process audio signals, data, recordings, etc., the input may or may not be the audio data itself and some ground truth data about the audio data. Then, once trained, the artificial intelligence system may receive some unknown audio data and output decision data about that unknown audio data. For example, the output decision data may relate to extracted sounds, stems, frequencies, etc., or to other decisions or AI decision observations of the input audio data.

[0056] The resulting data structures corresponding to the trained model can then be ported or distributed to other computer systems, which can then use the trained model. Once trained, the computer code in program memory, when executed by a processor, can receive an image as input, and based on the fact that the data structures represent the trained AI model being trained, the program code can process the input and output a decision about the nature of the input. In some implementations, the AI ​​model may comprise the program code and data structures that are entangled and not easily separated.

[0057] A model may be represented by program code and data, including a set of weights assigned to edges in a neural network graph, instructions on the neural network graph and how to interact with the graph, a mathematical representation such as a regression or classification model, and / or other data structures that may be publicly known in the art. A neural model (or neural network, or neural network model) can often be embodied as a data structure that shows or represents a set of connected nodes, often referred to as neurons, many of which may be data structures that mimic or simulate signal processing performed by biological neurons. Training may involve steps of updating parameters associated with each neuron, such as the weights and / or functions of the inputs and outputs of the neurons, and other neurons to which it is connected. In practice, a neural model can pass and process input data variables in some way to produce output variables to achieve a certain purpose, for example, to produce a binary classification of whether an input image or input dataset fits into a certain category. The training process may involve complex computations (e.g., computing gradient updates and then using the gradients to update parameters layer by layer). Training may be performed using some form of parallel processing.

[0058] Any of the audio source separation models described herein may also be referred to, at least in part, by several terms, including, for example, “audio source separation model,” “recurrent neural network,” “RNN,” “deep neural network,” “DNN,” “inference model,” or “neural network.”

[0059] Figures 1 and 2 illustrate an audio source separation system and process according to one or more implementations of the present disclosure. The audio processing system 100 includes a core machine learning system 110, a core model operation 130, and a modified recurrent neural network (RNN) class model 160. In the illustrated implementation, the core machine learning system 110 implements an RNN class audio source separation model (RNN-CASSM) 112, as depicted in a simplified representation in Figure 1. As illustrated, the RNN-CASSM 112 receives a signal input 114, which is input to a time-domain encoder 116. The signal input 114 may be a single-channel audio mixture, which may be received from a stored audio file accessed by the RNN-CASSM network, an audio input stream received from a separate system component, or another audio data source. The time-domain encoder 116 models the input audio signal in the time domain and estimates the audio mixture weights. In some implementations, the time-domain encoder 116 segments the audio signal input into separate waveform segments that are normalized for input to a 1D convolutional coder. The RNN class mask network 118 is configured to estimate a source mask for separating the audio source from the audio input mix. The mask is applied to the audio segment by the source separation component 120. The time-domain decoder 122 is configured to reconstruct the audio source, which is then available for output through the signal output 124.

[0060] In the exemplary implementation, the RNN-CASSM112 is modified for operation by the core model operation 130, which will be described herein. Referring to block 1.1, the audio source data is sampled using a 48 kHz sample rate in various implementations. Thus, the RNN-CASSM112 may be trained at frequencies higher than the audible frequency range to recognize distinct stems of audio in a lower frequency range. It has been observed that implementations of speech separation models trained at sample rates lower than 48 kHz produce lower quality audio separation in audio samples recorded using older equipment (e.g., 1960s Nagra® equipment). Various steps disclosed herein, such as training the source separation model, setting appropriate hyperparameters for its sample rate, and operating the signal processing pipeline at a 48 kHz sample rate, are performed at a 48 kHz sample rate. Other sampling rates, including oversampling, may also be used in other implementations in a manner consistent with the teachings of this disclosure. In block 1.2, the encoder / decoder framework (e.g., time-domain encoder 116 and time-domain decoder 122) is set to a step size of one sample (e.g., the input signal sample rate).

[0061] Referring to block 1.3, a modified RNN-CASSM 160 is generated, which extends beyond the source separation and noise reduction of conventional implementations. It has been observed that processed audio mixtures with trained RNN-CASSMs can sometimes result in undesirable click artifacts, harmonic artifacts, and broadband noise artifacts. To address this, the audio processing system 100 applies modifications to the RNN-CASSM network 112, as referred to in blocks 1.3a, 1.3b, and / or 1.3c, to reduce and / or avoid artifacts and provide other advantages. The same modifications may also be used for more transformative types of trained processing.

[0062] The model operations associated with the modified RNN-CASSM160 will be described in more detail here with reference to block 1.3ac. In block 1.3a, at least some of the encoder and decoder layers are removed and not used in the modified RNN-CASSM. It is observed that these removed layers may be redundant in the trained filter when the step size is one sample. In block 1.3b, the step of applying the mask (e.g., component 120) is also removed with respect to at least some part of the audio processing. Thus, the modified RNN-CASSM160 may perform audio source separation without applying a mask. The masking step potentially contributes to click artifacts, which are often present in the generated audio stem. To address this, the audio processing system may omit some or all of the masking and instead use the output of the RNN class mask network 118 in a more direct manner (e.g., train the RNN to output one or more isolated audio sources). In block 1.3c, the window function is applied to the overlapping summation step of the RNN class mask network (e.g., RNN network 162). When the model output has a linear harmonic sequence "banding" artifact at frequencies related to the overlapping summation function segment length, the audio processing system can smooth out hard edges when reconstructing the audio signal by unfolding the window function across each overlapping segment.

[0063] Referring to block 1.4, another core model operation 130 is the use of a separation strength parameter, which allows control over the strength of the separation mask applied to the input signal to produce separated sources. To provide direct control over the strength of the separation mask applied to the input signal to produce separated sources, a parameter is introduced during the model's forward pass that determines the degree to which the separation mask is applied. In one embodiment, the separation strength parameter is f(M) = M SThis can also be expressed as a function applied to the separation mask in a mask-type source separation model, where the mask M has a value [0, l] and s is the separation strength parameter. In this example, a value of s > 1.0 results in a lower mask value and fewer target sources when the mask is applied to the input mixture, while a value of s < 1.0 results in a higher mask value and a combination of target sources and complementary and noise components in the signal.

[0064] The separation strength parameter may be implemented as a helper function in an automated version of the Self-Iterative Training (SIPT) algorithm, which will be described in more detail below. It should be understood that RNN-CASSM112 may implement one or more of the core model operations 130 disclosed herein, including additional operations consistent with the teachings of this disclosure.

[0065] Referring to Figure 3-5, an exemplary process for training an RNN-CASSM network for audio source separation will be described here. The training process includes a plurality of labeled machine learning training datasets 310 configured to train a network for audio source separation as described herein. For example, in some implementations, the training dataset may include an audio mix and ground truth labels that identify the source class to be separated from the audio mix. In other implementations, the training dataset may include separated audio stems having audio artifacts (e.g., clicks) generated by the source separation process and / or one or more audio enhancements (e.g., reverb, filters), and ground truth labels that identify improved audio stems from which the identified audio artifacts and / or enhancements have been removed.

[0066] During operation, the network is trained by feeding it labeled audio samples. In various implementations, the network includes multiple neural network models that can be trained separately for specific source separation tasks (e.g., separating speech, separating foreground speech, separating drums, removing artifacts, etc.). Training includes a forward pass through the network to generate audio source separation data. Each audio sample is labeled with a "ground truth" that defines the expected output, which is compared to the generated audio source separation data. If the network mislabels an input audio sample, a backward pass through the network may be used to adjust the network's parameters to correct the misclassification. In various implementations, [ka] The output estimate, expressed as , is compared to the ground truth, expressed as Y, using regression losses such as an L1 loss function (e.g., minimum absolute deviation), an L2 loss function (e.g., least squares), a scale-invariant signal-to-distortion ratio (SISDR), a scale-dependent signal-to-distortion ratio (SDSDR), and / or other loss functions known in the art. After the network is trained, a validation dataset (e.g., a set of labeled audio samples not used in the training process) may then be used to measure the accuracy of the trained network. The trained RNN-CASSM network may then be implemented in the runtime environment to generate separate audio source signals from an audio input stream. The generated separate audio source signals may also be referred to as generated multiple audio stems. The generated multiple audio stems may correspond to one or more audio sources of the multiple audio sources in the audio input stream.

[0067] In various implementations, the modified RNN-CASSM network 160 is trained on a number of audio samples (e.g., thousands) that include audio samples representing multiple speakers, instruments, and other audio source information under various conditions (e.g., various noisy conditions and audio mixes). From the error between the separated audio source signals and the labeled ground truth associated with the audio samples from the training dataset, the deep learning model learns parameters that enable the model to separate the audio source signals. The RNN-CASM 120 and modified RNN-CASM 160 may also be referred to as the inference model 120 and the modified inference model 160 and / or the trained audio separation model 120 and the updated audio separation model 160.

[0068] Figure 3 is a schematic diagram of a machine learning dataset 310 and a training data loader 350, which may be relevant to the dataset and dataset manipulation during training by the audio processor, potentially providing improvements and / or useful functionality in output signal quality source isolation. In the illustrated implementation, an RNN-CASSM network is trained to isolate audio sources from a single-track recording of a music recording session, which includes a mixture of people singing / speaking, instruments being played, and various ambient noises.

[0069] Referring to Block 2.1, the training dataset used in the illustrated implementation includes a 48kHz speech dataset. For example, the 48kHz speech dataset may include identical utterances recorded simultaneously at various microphone distances (e.g., close microphones and farther microphones). In one test implementation, 85 different speakers, each with more than 20 minutes of utterances, were included in the 48kHz speech dataset. In various implementations, exemplary speech datasets may be created using a large number of speakers, such as 10, 50, 85, or more, from adult male and female speakers, with extended utterance periods of 10, 20 minutes, or more, and recordings at 48kHz or higher sampling rates. It should be understood that other dataset parameters may also be used to generate training datasets in accordance with the teachings of this disclosure.

[0070] Referring to Block 2.2, the training dataset further includes non-speech music and noise datasets, including segments of input audio such as digitized monaural audio recordings originally recorded on analog media. In some implementations, this dataset may also include segments of recorded music, non-vocal sounds, background noise, audio media artifacts from digitized audio recordings, and other audio data. Using this dataset, an audio processing system can more easily isolate the speaker's speech of interest from other speech, music, and background noise in the digitized recording. In some implementations, this may include the step of using manually collected segments of recordings that are manually noted as lacking speech or another audio source class, and labeling those segments accordingly.

[0071] Referring to Block 2.3, the dataset is generated and modified using a progressive self-repeating dataset generation process with a target unidentified mixture (e.g., a mixture unknown to the audio processing system). The generated dataset may include labeled datasets generated by processing unlabeled datasets (e.g., the target unidentified mixture to be source-separated) through one or more neural network models to generate the initial classification. This “roughly separated data” is then processed through a sweep step configured to select from the “roughly separated data” to retain the most useful “roughly separated data” based on a usefulness metric. For example, the performance of the training dataset may be measured by applying the validation dataset to a model trained with various training datasets containing the “roughly separated data” and determining, based on the calculated validation error, which data samples contribute to better performance and which contribute to poor performance. The usefulness metric may be implemented as a function that estimates a quality metric in the roughly separated data to identify “low-quality” tuned data that should be discarded before the next tune-up iteration. For example, a moving root-mean-square (RMS) window function may be calculated on the network's output to identify segments of output where the RMS metric lies above a calibrated threshold over a certain (or minimum) duration of the sample. This metric can be used, for example, to identify low-amplitude segments of coarsely separated source outputs that are more likely to produce artifacts. The threshold and minimum duration may be user-adjustable parameters that allow for adjustment of which data should be discarded.

[0072] The incremental self-repeating dataset generated using the unidentified target mixture may be generated using a self-repeating dataset generation loop 420, as illustrated in Figure 4. In various implementations, previous recordings containing sources to be separated that are still unidentified by the trained source separation model are less likely to be successfully separated by the model. Existing datasets may not be substantial enough to train a robust separation model for targeted sources, and there may be no opportunity to capture new recordings of sources. Instead of capturing new recordings of sources, additional training data may be manually labeled from isolated instances of sources in previous recordings. For example, isolated utterances from identified speakers, isolated audio of identified instruments recorded with similar equipment in similar environments, and / or other available audio segments that approximate the isolated sources may be added to the training data manually and / or automatically (based on metadata labeling of audio sources such as source identification, source class, and / or environment). This additional training data may be used to help fine-tune the model to improve its processing performance for previous recordings. However, this labeling process can involve a considerable amount of time and manual work, and there may not be sufficient isolated instances of the source in previous recordings to yield a sufficient amount of additional training data.

[0073] The illustrated generation elaboration tool can overcome these difficulties. In one method, a coarse general model 410 is trained on a general training dataset. The general training dataset may comprise labeled source audio data and labeled noise audio data. The general model 410 may be referred to as a general source separation model 410 or a trained audio source separation model 410. The training dataset may comprise multiple datasets, each comprising labeled audio samples configured to train the system to address the source separation problem. The multiple datasets may comprise a speech training dataset comprising multiple labeled speech samples and / or a non-speech training dataset comprising multiple labeled music and / or noise data samples. The available previous recordings, including an unidentified audio mixture to be separated, are then processed in process 422 using the general model to yield two labeled datasets of isolated audio (e.g., audio stems), namely, a coarse separated previous recording source dataset 424 and a coarse separated previous recording noise dataset 426.

[0074] Various implementations may also use other training datasets that provide a set of labeled audio samples selected to train the system to solve a specific problem (e.g., utterance vs. non-utterance). In some implementations, for example, the trained dataset may include (i) music vs. sound effects vs. Foley, (ii) a dataset of various instruments in a band, (iii) multiple human speakers separated from each other, (iv) a source from room reverberation, and / or (v) other training datasets. The results are then culled using a threshold metric (process 428) and audio windows below a selected root mean square (RMS) level, which may be user-selectable. In some implementations, the moving RMS may be calculated by segmenting the audio data into overlapping windows of equal duration and calculating the RMS for each window. The RMS level may be referred to as a utility metric or a quality metric, and the RMS level may be one of the alternative utility metric or quality metric options. The quality metric may be calculated based on multiple associated audio stems.

[0075] Next, the new model is trained in process 430 using a culled self-repeating dataset (for example, culled results added to the audio training dataset; the culled self-repeating dataset is also referred to as the culled dynamically evolving dataset) to train and generate an improved model 432, improving its performance when processing recordings. The improved model 432 may be configured to reprocess the audio input stream and generate multiple improved audio stems. This improved model is an update to the trained audio source isolation model. This improved model may be referred to as the updated audio source isolation model 432.

[0076] In some implementations, the audio training dataset may be curated during the iterative fine-tuning process to remove data irrelevant to the target input mixture. For example, the input mixture may identify / classify various sources within it, leaving some other source categories that are not identified / classified and / or otherwise irrelevant to the source separation task. Training data associated with these "irrelevant" source categories (e.g., categories not found in the target mixture, categories identified by the user as irrelevant to the source separation task, and / or other irrelevant source categories as defined by other criteria) may be culled from the audio training dataset, allowing the training dataset to become increasingly specific to the content of the target input mixture.

[0077] Process 420 is repeated iteratively, each time improving the separation quality of the model (e.g., fine-tuning to improve the accuracy and / or quality of source separation). Depending on subsequent re-iterations, process 420 may use further RMS levels, thereby exceeding the previous RMS level. This process enables more automated refinement of the initial general source isolation or separation model. Looping at various stages has been observed to show superior improvement over larger related mixtures. The general model 410 and the improved model 432 may also be referred to as the inference model 410 and the modified inference model 432 and / or the trained audio separation model 410 and the updated audio separation model 432.

[0078] Improvements in separation quality (e.g., audio fidelity) can be measured by the system and / or evaluated by a user who provides feedback to the system through a user interface and / or oversees one or more steps in process 420. In some implementations, process 420 may use a combination of an algorithm for calculating a mean opinion score (MOS) and / or user evaluation to estimate separation quality. For example, the algorithm can estimate the amount of artifacting that occurred during source separation operations, which in turn relates to the overall quality of the network's output. In some implementations, estimating the amount of artifacting that occurred during source separation operations includes a step of feeding the separated sources through a neural network model trained to separate audio artifacts from the signal, allowing for the measurement of the presence and / or intensity of such audio artifacts. The intensity of the audio artifacts is determined in each iteration and can be tracked between iterations to fine-tune the model. In some implementations, the iterative process continues until the estimated separation quality across the iterations no longer improves and / or the estimated separation quality meets a predetermined quality threshold of one or more.

[0079] Referring to blocks 2.4 and 2.4a, the machine learning training data loader 350 is configured to match the sonic quality of the target unidentified mixture during training (e.g., perceived distance from the microphone, filtering, reverberation, echo, nonlinear distortion, spectral distribution, and / or other measurable audio quality). The challenge in training an effective supervised source separation model is that the dataset examples must be curated to match as closely as possible to the quality of the sources in the target mixture, ideally. For example, speakers should be isolated from the mixture in which they are speaking into a microphone in a reverberant hall, and the recording may be compared to how it would sound if the same speaker were recorded speaking directly into a microphone in a neutral, reverberant space, such as where a high-quality speaker dataset might be recorded, if the recording was captured from some distance by a recording device in an audience. The goal, then, is to add extensions to the high-quality speech dataset samples during training, which generally result in lower deviations from the target input mixture. In this embodiment, a “reverberation” extension may be added to simulate that of a hall, a “nonlinear distortion” extension to simulate the speaker’s voice amplified by the acoustic system, and a “filter” extension to simulate the distance of the acoustic system from the recording device.

[0080] In various implementations, the solution involves a hierarchical mix bus schema, including pre- and post-mixture extension modules. Creating the ideal target mixture would be cumbersome or impractical to manually create each time a new type of mixture is required for training. The hierarchical mix bus schema allows for the easy definition of arbitrarily complex randomized "sources" and "noise" during supervised source separation training. The data loader roughly matches quality such as filtering, reverberation, relative signal level training, and other audio qualities for improved source separation results. The machine learning data loader uses a hierarchical schema that allows for the easy definition of dynamically generated "source" and "noise" mixtures from source data while training the model. The mix bus allows for arbitrary extensions such as reverb or filters with accompanying randomization parameters. By using a well-classified dataset medium as raw material, this allows for the easy setup of a training dataset that mimics the desired source separation target mixture.

[0081] An exemplary simplified schema representation 550 is illustrated in Figure 5. The training mixture schema includes separate choices for source and noise, and includes criteria such as dB range, probability associated with source determination, room impulse response, filter, and other criteria.

[0082] Referring to block 2.4b, the data loader further provides pre / post-mixture augmentations, including filters, nonlinear functions, and convolutions, which are applied to the target mixture during training. In various implementations, the relevant augments are identified and added to the training dataset, and the isolated audio stems are post-processed (e.g., using the pipelines described herein). The source separation model can be trained with an additional objective of transforming the isolated sources using the augments. In some implementations, the system can be trained to strictly isolate sources where they might be audible in the input mixture. In some implementations, the system may further be trained to improve the quality of some of the isolated sources by applying appropriate augments, e.g., augments that result in the smallest deviation from the target mixture when using the available training dataset. Deviations and appropriate augments may be estimated algorithmically and / or by user evaluation. For example, speech in the input mixture may be filtered and difficult to understand due to being recorded behind a closed door. In this embodiment, the isolated sources may be augmented (e.g., degrading the input speech dataset to generally match the sound of speech behind a closed door). However, the isolated source outputs of the target during training are not augmented in this embodiment (e.g., augmented speech input versus corresponding high-quality speech target output dataset during training), and therefore the network is trained to approximate this same transformation by post-processing augmentation.

[0083] In some implementations, the target source may be a transformed version of its current representation within a mixture. For example, there may be a need to restore a bandwidth-limited recording to a more complete frequency spectrum, or to isolate and increase the proximity fidelity of obscured background speakers. These needs may be left to the transformative model itself to be resolved either by user determination or automatically based on input deviations from the target output training set. For example, when a model is trained to output high-quality near utterances with little reverberation, using a speech dataset that has been randomly augmented during training, inputting a mixture containing such high-quality near utterances with little reverberation may tend to result in minimal changes to those inputs. However, inputting a mixture containing utterances that deviate from these qualities may tend to transform those input speech mixtures to resemble high-quality near utterances with little reverberation.

[0084] The extension module can be used by the data loader to generate training examples consisting of an extended source as input, along with an alternatively extended version of the same source as the target output. This enables transformative examples where the target source can be represented in an alternatively extended context when it is part of the training mixture. When used while training the modified RNN-CASSM, this allows the audio processing system to learn behaviors such as "filtering removal," "reverberation removal," and deeper recovery of a highly obscured target source.

[0085] The illustrated implementation example 500 includes filtering removal, which includes the step of (i) expanding the mixture using a filter; (ii) de-reverberation, which includes the step of expanding the mixture using a reverb; (iii) background speaker recovery, which includes the step of expanding the source using a filter and a reverb; (iv) distortion repair, which includes the step of expanding the mixture using distortion; and (v) gap repair, which includes the step of expanding the mixture using a gap.

[0086] Referring to Figures 6 and 7, exemplary implementations of machine learning training method 600 will be described here. In these embodiments, the machine learning training method is described in relation to a method for training for improvements and / or useful functionality in output signal quality and source separation. Referring to block 3.1, the first machine learning training method includes a step of upscaling the trained network sample rate (e.g., from 24 kHz to 48 kHz). Due to limitations such as time and computational resources, the model may be trained at 24 kHz using a 24 kHz dataset, but with limitations on output quality. An upscaling process initiated for a model trained at 24 kHz can provide functionality at 48 kHz. One exemplary process includes a step of preserving the inner block of the learned parameters of the masking network while discarding the encoder / decoder layers and those directly connected to them. In other words, only the inner separation layer is ported into a newly initialized model with a 48 kHz encoder / decoder. The untrained 48 kHz encoder / decoder and their directly connected layers are then fine-tuned using a 48 kHz dataset while the inherited network remains frozen. This is done until an acceptable validation / loss value (e.g., L1, L2, SISDR, SDSDR, or other loss calculation) is confirmed again during training / validation, here indicating the fitted inherited layers. For example, an acceptable validation / loss value may be determined by comparing it to a given threshold loss value, by observing the trend toward a value confirmed in a previous training session compared to the model performance, or by other approaches. During training, the loss value ideally tends toward minimization; however, in practice, the loss value can also be useful in signaling significant problems during training, such as when the loss value begins to trend away from minimization. It has also been observed that the loss value may not improve during training, but the model's performance, such as that measured by the quality of source separation, can still be improved by continuing training.

[0087] Finally, fine-tuning training continues across all layers, allowing the model to further develop at 48kHz, with the final result being a well-functioning 48kHz model. In some implementations, the system is trained to operate more quickly at higher signal processing sample rates by employing a two-step training process: firstly, the untrained layers are trained while the parameters of the inherited layers are frozen; and secondly, the entire model is then fine-tuned until its performance matches or exceeds that of the lower-sample-rate model. This process can be carried out over many iterations, including, but is not limited to:

[0088] a) Train at 6kHz

[0089] b) Upscale to 12kHz

[0090] c) Upscale to 24kHz

[0091] d) Upscale to 48kHz

[0092] Referring to Block 3.2, multi-source mixtures are used to improve the performance of speech isolation models (e.g., source = foreground, background, and far-field speech, noise, and music mixture). Utterances initially trained with a single speech-versus-noise / music mixture may not perform well, and the processed results may have difficulty consistently extracting from the original source medium and may suffer from substantial artifacting. Instead of training with a single speech-versus-noise mixture, an audio processing system may provide substantially improved results by using multiple layered speech sources to simulate fluctuations in proximity, such as foreground and background speech. This approach can also be applied to instruments, e.g., multiple overlaid guitar samples in the mixture instead of just one sample at a time. In various implementations, training samples with various layering scenarios may be selected to match and / or approximate an unidentified audio mixture (e.g., based on user input, identified source classes, and / or analysis of the unidentified audio mixture during iterative training).

[0093] Training mixture example 700 includes a speech isolation training mixture 702 which includes a mixture of source and noise 704. Source mixture 706 may include a mixture of foreground utterances 708, background utterances 710, and distant utterances 712 in an expanded randomized combination. Noise mixture 714 includes a mixture of instrument 716, room tone 718, and hiss 720 in an expanded randomized combination.

[0094] Referring to Figure 8-10, the implementation of the machine learning process 800 will be described here in relation to a method of processing using a machine learning model that contributes to improvements and / or useful functionality in output signal quality and source separation (for example, as discussed above in this specification). Referring to Block 4.1, the machine learning process may include the interpolation of the sum of separated sources as an additional output. In other implementations, the model outputs speech and discarded music / noise. The interpolated output may subsequently be used in various processes to further process / separate the sources remaining in the interpolated output.

[0095] Referring to Block 4.2, the machine learning post-processing model cleans up artifacts introduced by the machine learning process, such as clicks, harmonic distortion, ghosting, and broadband noise (artifacts may be determined as discussed above, for example, with reference to Figure 4). The source separation model being trained may exhibit artifacting such as clicks, harmonic distortion, broadband noise, and "ghosting," where sounds are partially separated between the target interpolated outputs. For these outputs to be used in the context of a high-quality soundtrack, a laborious cleanup would typically need to be attempted using conventional audio restoration software. Attempting to do so still may result in undesirable quality in the restored audio. This can be addressed by post-processing the audio processed with a model trained on a dataset of processing artifacts. The processing artifact dataset may be generated by the problematic model itself.

[0096] The post-processing model 910, once trained, can be reused for all similar models. In the illustrated implementation, the input mixture 950 is processed using a general model 952, which generates a source-separated output with machine learning artifacts (step 954). The post-processing step 956 removes the artifacts and generates an improved output 960. The post-processing model 910 includes the step of generating a dataset 914 of isolated machine learning artifacts (step 912). The machine learning artifacts may include clicks, ghosting, broadband noise, harmonic distortion, and other artifacts. The isolated machine learning artifacts 912 are used in step 916 to train a model to remove the artifacts.

[0097] Referring to Block 4.3, in some implementations, a user-inducing self-repeating processing / training approach may be used. The user induces and contributes to fine-tuning the pre-trained model throughout the processing / editing / training loop, which can then be used to produce better source separation results than what might have been possible from a generic model. Model fine-tuning ability can be left to the user to solve source separations that the pre-trained model may not be able to solve. Processing unidentified mixtures with source separation models is not always successful, usually due to a lack of sufficient training data. In one solution, the audio processing system may induce and contribute to the user fine-tuning the pre-trained input, which can then be used to produce better results. In the exemplary method, i) the user processes the input medium, ii) is given the opportunity to assess the output, or has the option to have this assessment carried out by an algorithm that measures threshold parameters for several metrics (e.g., measured using a moving RMS window and / or other measurements as discussed above), iii) if the output is deemed acceptable, processing ends here; otherwise, iv) is given the opportunity to work with the incomplete output using temporal and / or spectral editing and / or culling / extension algorithms. Essentially, the user selects the segment of the output that will be most useful during the imminent steps. In step v), the medium is considered here in terms of inclusion into the training dataset, in step vi), the user is also given the opportunity to add their own auxiliary dataset, in step vii), the model is fine-tuned and trained according to the user's hyperparameter preferences, in step viiii), the model's performance is validated to confirm improved results from previous iterations, and then in step ix), the process is repeated. In various implementations, hyperparameters related to fine-tuning training may include parameters such as training segment length, epoch duration, training scheduler type and parameters, optimizer type and parameters, and / or other hyperparameters.The hyperparameters are initially based on a predetermined set of values, and may then be modified by the user for fine-tuning training.

[0098] An exemplary user-inducing self-repeating process 1000 is illustrated in Figure 10. Process 1000 begins with a general model 1002, which can be implemented as a pre-trained model as discussed above. An input mixture 1004, such as a single-track audio signal or multiple single-track audio signals with an unidentified source mixture, is processed through the general model 1002 in step 1006 to generate separated audio signals from the mixture, which include a machine learning-separated source signal 1008 and a machine learning-separated noise signal 1010. In step 1012, the results are evaluated to confirm that the separated audio sources have sufficient quality (for example, by comparing them to estimated MOS and threshold and / or other quality measurements as described above). If the results are determined to be good, the separated sources are output in step 1014.

[0099] If it is determined that the separated audio source signals require further improvement, one or more of the outputs are prepared in step 1016 for inclusion in the training dataset for fine-tuning. In the automated fine-tuning system 1018, the machine learning-separated sources 1008 and noise 1010 are used directly as the fine-tuning dataset 1034 (step 1022). The fine-tuning dataset is optionally culled in step 1036 based on a user-selected threshold for the source separation metric (e.g., moving RMS window compared to thresholds and / or other separation metrics as described above). Training is then performed in step 1038 to fine-tune the model. The fine-tuned model 1032 is then applied in step 1006 to the input mixture.

[0100] In the user-guided fine-tuning system 1020, the user may select portions of audio clips to include in or omit from the fine-tuning dataset (step 1024 - temporal editing). The user may also select portions of clips to include in / omit from the fine-tuning dataset and frequency / time window selections (step 1026 - spectral editing). In some implementations, the user may provide additional audio clips to augment the fine-tuning dataset (step 1028 - adding to dataset). In some implementations, the user may provide equalization, reverb, distortion, and / or other augmentation settings to fine-tune the dataset (step 1030 - augmentation). After the fine-tuning dataset 1034 has been updated, the training process continues through steps 1036-1038 to produce the fine-tuned model 1032.

[0101] Referring to Block 4.4, an animated visual representation of the model fine-tuning progress may be implemented. While the user fine-tunes and trains the model to resolve source separation for a particular media clip, the model's progress output is displayed to help demonstrate the model's performance in order to help guide the user's decision-making. In some implementations, for example, the interface may be displayed in a window with associated tool icons that displays a periodically updated spectrogram animation representation of the estimated outputs, such as those computed by the fine-tuned model, which are periodically tested while this is being fine-tuned and trained. This interface can allow the user to visually assess the extent to which the model is currently performing well within various regions of the input mixture. The interface may also facilitate user interaction, such as allowing the user to experiment with time / frequency selections of these estimated outputs based on their interaction with the spectrogram window.

[0102] Referring to Block 4.5, user-inducing extensions for fine-tuning training may be implemented. To improve the results when enhancing / separating targeted recordings with specific characteristics such as reverberation, filtering, and nonlinear distortion, the audio processing system may present the user with tools to guide and control the underlying algorithms for contributing to the selection of extensions such as reverberation, filtering, nonlinear distortion, and / or noise during step 1020 of the loop described in Block 4.3. The user can guide and contribute to the selection of extensions and / or extension parameters used during the fine-tuning of the pre-trained model throughout the process of the processing / editing / training loop. The extensions include user-controllable randomization settings for each parameter (e.g., values ​​affecting various aspects of the extension such as the intensity, density, modulation, and / or decay of the reverberation extension) to help generalize or narrow the targeted behavior after fine-tuning. This allows for increased control that the user has over fine-tuning training, which allows them to specifically match, for example, filtered sounds / reverberation in the targeted recording. When referring to a random process or randomization, it may be sufficient to have a pseudo-random process or an arbitrary selection process. In some implementations, the extension parameters are automatically matched to the input source mixture using one or more algorithms. In some implementations, automatic matching is used to achieve a rough match as a starting point. For example, spectral analysis of the target input mixture may be combined with analysis of a random dataset sample to yield a set of frequency band deviation scores, which can be used to minimize deviations between the randomly extended dataset sample and the target input mixture by adjusting various parameters of the dataset extension filter based on the values ​​of the frequency band deviation scores.

[0103] Referring to Figures 11A and 11B, exemplary implementations of the machine learning application 1100 will be described here according to one or more implementations. The audio processing systems and methods disclosed herein may be used in conjunction with other audio processing applications that contribute to improvements and / or useful functionality for audio post-production editing workflows.

[0104] Referring to Block 5.1, a plugin such as an Avid Audio Extension (AAX) plugin is provided, which is hosted in a digital audio workstation (DAW) (for example, a digital audio workstation sold under the trademark name PRO TOOLS may be used in one or more implementations), enabling the user to submit audio clips from a standalone application, where they may subsequently be processed using a machine learning model. The plugin can return an arbitrary number of stem splits to the DAW environment.

[0105] Referring to Block 5.2, implementations of the present disclosure are also used in applications (e.g., applications marketed as JAM LAB with JAM CONNECT or similar applications) to load / receive media, which are then processed by user-selected machine learning recipes. In some implementations, multiple client machines 1 102A-C and processing nodes 1 106A-D, having access to client software (e.g., JAM LAB with JAM CONNECT), are configured to access a task manager / database 1104 for access to both client software and machine learning (ML) applications 1100 as disclosed herein.

[0106] Referring to Figure 12, an exemplary processing flow will be described here. A system running a digital audio workstation 1202 (e.g., PRO TOOLS with JAM CONNECT) is configured to send audio clips to a client application 1208 (e.g., including JAM LAB). Through the client application, a single model may be selected from a list of categories 1210 to process / separate the audio source from the audio clip. The type of stem is also selected in step 1212 to form a multi-model recipe. The audio clips and recipes are sent (in step 1214) to a task manager / database 1216, which in step 1218 manages and distributes recipe processing across the user's available processing nodes. In step 1206, the client application receives and returns the processed audio clips labeled with the stem and / or model name. In some implementations, the client application 1208 may also facilitate the selection of stems to form a multi-model recipe 1212 as disclosed herein.

[0107] An implementation of a machine learning multi-model recipe processing system will be described here with reference to Figures 13A-E. In some implementations, the user may desire to separate a targeted medium into a set of source classes / stems in a single step using a selection of source separation models in a specific order and hierarchical combination. Sequential / branching source separation recipe schemas are implemented to process targeted medium using one or more source separation models in a sequential / branching structured order and separate the targeted medium into a user-selected set of source classes or stems.

[0108] In implementation 1300 of Figure 13A, each step of the recipe represents a source separation model or combination of models that targets a specific source class. The recipe is defined according to a recipe schema for performing steps that yield the user's desired stem output and includes a well-trained model. As illustrated in the exemplary implementation, the user may choose to separate hiss first, then voice (then separated into vocals and other utterances), drums (then separated into kick, snare, and other percussion), organ, piano, bass, and other processing. Thus, by processing the targeted medium through the pipeline defined by the recipe, the outputs at various steps are collected and ultimately the targeted medium may be separated into a user-selected source class or set of stems.

[0109] Referring to Figure 13B, an example of a speech / drum and other sequential processing pipeline 1320 is illustrated. In this model, the input mixture is first processed to extract the speech, and the completion includes the drum and other sounds in the mixture. The drum model extracts the drum, and the completion includes the other sounds. In this implementation, the output includes the speech, the drum, and other stems. The order in which the models are applied when separating source classes using a sequential / branching separation system may be optimized for higher quality by using an algorithm that assesses the optimal processing order. An exemplary implementation of an optimized processing method 1340 is illustrated in Figure 13C. For example, an input mixture model having components A and B may be configured to isolate class A, and then the remainder B. This may yield different results if the processing order is reversed (e.g., isolating B, and then A). The optimized processing method 1340 may operate on an input mixture of A+B by separating A and B in both orders, comparing the results, and selecting the order that yields the best results (e.g., the result with less error in the separated stems). In various implementations, the optimized processing method 1340 may operate manually, automatically, and / or in a hybrid approach. For example, the estimated optimal order may be pre-established by using a set of ground truth test samples and then creating a test mixture that can be separated using various sortings of stem orders, and thus the estimated best-performing model processing order may be established by using an error function that compares the outputs of these tests against the ground truth test samples.

[0110] An exemplary pipeline 1360 for improving the output fidelity of a sequential / branching source separation system (e.g., as described above herein) will be described here with reference to Figure 13D. For example, output fidelity may be measured using a MOS algorithm, which is capable of measuring the deviation of a stem's output from a particular labeled dataset, such as a collection of speech samples from an individual. In some implementations, such an algorithm may be implemented as a neural network pre-trained for either classifying sources or measuring the deviation of sources from a given dataset.

[0111] The processing recipe (see 5.2.1) may also include a post-processing step after one or more outputs in the pipeline. The post-processing step may include any type of digital signal processing filter / algorithm to remove any signal artifacts / noise that may have been introduced by the preceding steps. The pipeline 1360 uses a model specifically trained to remove artifacts (for example, as described in step 4.2 of this specification) and, in particular, due to the sequential nature of the recipe processing pipeline, it yields a significantly improved overall result.

[0112] Referring to Figure 13E, exemplary implementation 1380 combines models to isolate a specific source class. Sequential / branching isolation processing recipes (see 5.2.1) may include a step in which one or more models are used in combination to extract from a mixture source classes that could not otherwise be fully extracted by a single model trained to target that source class. In the illustrated embodiment, the drum is separated twice, then summed and presented as a single “drum” output stem, accompanied by a complementary “other” stem.

[0113] An exemplary audio processing system 1400 for implementing the systems and methods disclosed herein will be described here with reference to Figure 14. The audio processing system 1400 includes a logic device 1402, a memory 1404, a communication component 1422, a display 1418, a user interface 1420, and a data storage device 1430.

[0114] The logic device 1402 may include, for example, a microprocessor, a single-core processor, a multi-core processor, a microcontroller, a programmable logic device configured to perform processing operations, a DSP device, one or more memories for storing executable instructions (e.g., software, firmware, or other instructions), a graphics processing unit, and / or any other suitable combination of processing devices and / or memories configured to execute instructions for performing any of the various operations described herein. The logic device 1402 is adapted to interface with and communicate with various components of the audio processing system 1400, including memory 1404, a communication component 1422, a display 1418, a user interface 1420, and a data storage device 1430.

[0115] The communication component 1422 may include wired and wireless communication interfaces to facilitate communication with a network or remote system. The wired communication interface may be implemented as one or more physical network or device connection interfaces such as cables or other wired communication interfaces. The wireless communication interface may be implemented as one or more Wi-Fi, Bluetooth®, cellular, infrared, radio waves, and / or other types of network interfaces for wireless communication. The communication component 1422 may include antennas for wireless communication during operation.

[0116] The display 1418 may include an image display device (e.g., a liquid crystal display (LCD)) or various other types of commonly known video displays or monitors. The user interface 1420 may, in various implementations, include user input and / or interface devices such as a keyboard, a control panel unit, a graphical user interface, or other user input / output devices. The display 1418 may operate as both a user input device and a display device, for example, a touchscreen device adapted to receive input signals from a user touching different parts of the display screen.

[0117] Memory 1404 stores program instructions for execution by logic device 1402, including program logic for implementing the systems and methods disclosed herein, including, but not limited to, an audio source isolation tool 1406, core model operations 1408, machine learning training 1410, a trained audio isolation model 1412, an audio processing application 1414, and self-repeating processing / training logic 1416. Data used by the audio processing system 1400 is stored in memory 1404 and / or may be stored in data storage device 1430, and may include a machine learning utterance dataset 1432, a machine learning music / noise dataset 1434, an audio stem 1436, an audio mixture 1438, and / or other data.

[0118] In some implementations, one or more processes may be implemented through a remote processing system, such as a cloud platform, which may be implemented as an audio processing system 1400 as described herein.

[0119] Figure 15 illustrates an exemplary neural network that may be used in one or more implementations of Figures 1-14, including various RNNs and models as described herein. The neural network 1500 may be implemented as a recurrent neural network, a deep neural network, a convolutional neural network, or any other suitable neural network that receives a labeled training dataset 1510 to generate an audio output 1512 (e.g., one or more audio stems) for each input audio sample. In various implementations, the labeled training dataset 1510 may include various audio samples and training mixtures as described herein, such as a training dataset (Figure 3), a self-repeating training dataset or a culled dataset (Figure 4), a training mixture and dataset as described herein according to the training method described herein (Figure 5-13E), or, as appropriate, other training datasets.

[0120] The training process for generating a trained neural network model includes a forward pass through a neural network 1500 to generate an audio stem or other desired audio output 1512. Each data sample is labeled with the desired output of the neural network 1500, which is compared to the audio output 1512. In some implementations, a cost function may be applied to quantify the error in the audio output 1512, and a backward pass through the neural network 1500 may then be used to adjust the neural network coefficients to minimize the output error.

[0121] The trained neural network 1500 may then be tested for accuracy using a subset of the labeled training data 1510 reserved for validation. The trained neural network 1500 may then be implemented as a model in the runtime environment to perform audio source separation as described herein.

[0122] In various implementations, the neural network 1500 uses the input layer 1520 to process input data (e.g., audio samples). In some embodiments, the input data may correspond to audio samples and / or audio inputs as described herein.

[0123] The input layer 1520 includes multiple neurons used to adjust the input audio data for input to the neural network 1500, which may include feature extraction, scaling, sampling rate conversion, and / or equivalent. Each neuron in the input layer 1520 generates an output, which is fed into the input of one or more hidden layers 1530. The hidden layer 1530 includes multiple neurons that process the output from the input layer 1520. In some embodiments, each neuron in the hidden layer 1530 generates an output, which is collectively then propagated through additional hidden layers, which include multiple neurons that process the output from the preceding hidden layers. The output of the hidden layer 1530 is fed into the output layer 1540. The output layer 1540 includes one or more neurons used to adjust the output from the output layer 1540 and generate a desired output. Please understand that the Neural Network 1500 architecture is merely representative, and other architectures are also possible, including neural networks with only one hidden layer, neural networks without input and / or output layers, neural networks with recurrent layers, and / or equivalents.

[0124] In some embodiments, the input layer 1520, the hidden layer 1530, and / or the output layer 1540 each contain one or more neurons. In some embodiments, the input layer 1520, the hidden layer 1530, and / or the output layer 1540 may each contain the same or different number of neurons. In some embodiments, each neuron takes a combination of its inputs x (e.g., a weighted sum using a trainable weighting matrix W), adds an optional trainable bias b, applies an activation function f, and generates an output α as shown by the equation α = f(Wx + b). In some embodiments, the activation function f may be a linear activation function, an activation function with upper and / or lower bounds, a log-sigmoid function, a hyperbolic tangent function, a rectified linear unit function, and / or equivalents. In some embodiments, each neuron may have the same or different activation functions.

[0125] In some embodiments, the neural network 1500 may be trained using supervised learning, where the combination of training data includes combinations of input data and ground truth (e.g., expected) output data. The difference between the generated audio output 1512 and the ground truth output data (e.g., markers) is fed back into the neural network 1500 to correct various trainable weights and biases. In some embodiments, the difference may be fed back using a backpropagation technique and / or equivalent using a stochastic gradient descent algorithm. In some embodiments, a large set of training data combinations may be presented to the neural network 1500 multiple times until the overall cost function (e.g., mean squared error based on the difference of each training combination) converges to an acceptable level.

[0126] An example implementation is described below.

[0127] 1. An audio processing system comprising a deep neural network (DNN) trained to separate one or more audio source signals from a single-track audio mixture.

[0128] 2. The audio processing system according to Embodiment 1, wherein the DNN is configured to receive a signal input and generate a signal output without time-domain encoding and / or time-domain decoding.

[0129] 3. The audio processing system according to Example 1-2, wherein the DNN is configured to apply a window function.

[0130] 4. The audio processing system according to Examples 1-3, wherein the DNN performs a duplicate addition process to smooth out banding artifacts.

[0131] 5. Audio source separation is performed without applying a mask in the audio processing system described in Examples 1-4.

[0132] 6. The audio processing system described in Examples 1-5, wherein the DNN model is trained using a 48 kHz sample rate.

[0133] 7. The audio processing system according to Examples 1-6, wherein the signal processing pipeline operates at 48 kHz.

[0134] 8. The audio processing system according to Examples 1-7, further comprising a separation strength parameter for controlling the intensity of the separation process applied to the input audio signal.

[0135] 9. The audio processing system according to Examples 1-8, further comprising a speech training dataset comprising multiple labeled speech samples.

[0136] 10. The audio processing system according to Examples 1-9, further comprising a non-speech training dataset comprising multiple labeled music and / or noise data samples.

[0137] 11. The audio processing system according to Examples 1-10, further comprising a dataset generation module configured to generate labeled audio samples for use in training a DNN model.

[0138] 12. The data set generation module is a self-repeating data set generator, as described in Example 1-11 of the audio processing system.

[0139] 13. The audio processing system according to Example 1-12, wherein the dataset generation module is configured to generate labeled audio samples from an input audio mixture and / or an audio source stem output from a DNN.

[0140] 14. The audio processing system according to Examples 1-13, further comprising a data loader configured to apply pre / post mixture expansion.

[0141] 15. The audio processing system according to Examples 1-14, wherein the DNN is trained at frequencies higher than the audible frequency range in order to recognize separate stems of audio within a lower audible frequency range.

[0142] 16. The audio processing system according to Examples 1-15, wherein the data loader is configured to apply extensions such as reverb, filters, and stochastic parameters.

[0143] 17. The audio processing system according to Examples 1-16, further configured to match the sound quality of a target unidentified mixture during training based on relative signal levels.

[0144] 18. Exemplary methods,

[0145] The process involves processing audio input data and generating source-separated stems using a trained inference model trained for source separation, and

[0146] The steps include generating a speech dataset from a source separation stem,

[0147] The steps include generating a noise dataset from a source separation stem,

[0148] The steps include training an inference model to generate an updated inference model using speech and noise datasets, Methods that include...

[0149] 19. The method according to Example 18, further comprising the step of processing audio input data using an updated inference model.

[0150] 20. The method according to Examples 18-19, further comprising the step of iteratively updating the updated inference model.

[0151] 21. The method according to Examples 19-20, wherein the training dataset is curated to include samples that approximate the audio source.

[0152] 22. The method according to Examples 19-21, further comprising a hierarchical mix bus schema.

[0153] 23. The inference model is trained using a multi-speech source mixture, as described in Examples 19-22.

[0154] 24. The method according to Examples 19-23, wherein the inference model is trained using foreground audio, background audio, and / or far audio.

[0155] 25. The method according to Examples 19-24, wherein the inference model is trained at a first sample rate and upscaled to a higher sample rate.

[0156] 26. The method according to Examples 19-25, further comprising the step of post-processing the separated audio source stems to remove artifacts introduced by the source separation process.

[0157] 27. The source separation stem is the method according to Examples 19-26, comprising a separated source signal and the remaining complementary signal.

[0158] 28. Artifacts introduced during the source separation process include clicks, harmonic distortion, ghosting, and / or broadband noise, as described in Examples 19-27.

[0159] 29. The fine-tuning process is as described in Examples 19-28, including a user-inducing self-repeating process.

[0160] 30. The method according to Examples 19-29, further comprising a step to facilitate user-guided extension for fine-tuning of training.

[0161] 31. A system comprising: an audio input configured to receive an audio input stream comprising a mixture of audio signals generated from a plurality of audio sources; a trained audio source separation model configured to receive the audio input stream and generate a plurality of generated audio stems, wherein the generated audio stems correspond to one or more of the plurality of audio sources; and a self-repeating training system configured to at least partially update the trained audio source separation model to an updated audio source separation model based on the generated audio stems, wherein the updated audio source separation model is configured to reprocess the audio input stream and generate a plurality of improved audio stems.

[0162] 32. The system according to Example 31, comprising a neural network, wherein the audio input stream comprises one or more single-track audio mixtures, and the trained audio source separation model is trained to separate one or more audio source signals from one or more single-track audio mixtures.

[0163] 33. The system according to Examples 31-32, wherein the neural network is configured to perform audio source separation without applying a mask.

[0164] 34. The system according to Examples 31-33, further comprising a training dataset comprising labeled source audio data and labeled noise audio data, wherein a trained audio source separation model is trained to generate a general source separation model using the training dataset.

[0165] 35. The system described in Examples 31-34, wherein at least a subset of the generated audio stems are culled based on a threshold metric and added to a training dataset to form a culled dynamically evolving dataset, which is used to train an updated audio source separation model.

[0166] 36. The self-repeating training system is further configured to calculate a first quality metric associated with multiple generated audio stems, the first quality metric providing a first performance measure of the trained audio source separation model, and the self-repeating training system is further configured to calculate a second quality metric associated with improved audio stems, the second quality metric providing a second performance measure of the updated audio source separation model, the second quality metric being superior to the first quality metric, as described in Examples 31-35.

[0167] 37. The system according to Examples 31-36, wherein the trained audio source separation model comprises multiple datasets, each comprising labeled audio samples, which are trained using a training dataset, and each of the multiple datasets is configured to train the system to address a source separation problem.

[0168] 38. The system according to Examples 31-37, wherein the multiple datasets include a speech training dataset comprising multiple labeled speech samples and / or a non-speech training dataset comprising multiple labeled music and / or noise data samples.

[0169] 39. The self-repeating training system further comprises a self-repeating dataset generation module configured to generate labeled audio samples from a plurality of generated audio stems, as described in Examples 31-38.

[0170] 40. The system described in Examples 31-39, in which multiple improved audio stems are generated using a hierarchical branching sequence that includes a step of separating the source signal and the remaining complementary signal.

[0171] 41. A method comprising: receiving an audio input stream comprising a mixture of audio signals generated from a plurality of audio sources; generating a plurality of generated audio stems corresponding to one or more of the plurality of audio sources using a trained audio source separation model configured to receive the audio input stream; updating the trained audio source separation model to an updated audio source separation model, at least partially based on the generated audio stems, using a self-repeating training process; and reprocessing the audio input stream using the updated audio source separation model to generate a plurality of improved audio stems.

[0172] 42. The method according to Example 41, wherein the audio input stream comprises one or more single-track audio mixtures, and the trained audio source separation model comprises a neural network trained to separate one or more audio source signals from one or more single-track audio mixtures.

[0173] 43. The method according to Examples 41-42, wherein the neural network is configured to perform audio source separation without applying a mask.

[0174] 44. The method according to Examples 41-43, further comprising the steps of providing a training dataset comprising labeled source audio data and labeled noise audio data, and training an audio source separation model trained to generate a general source separation model using the training dataset.

[0175] 45. The method according to Examples 41-44, further comprising the steps of: adding at least a subset of generated audio stems to a training dataset to generate a dynamically evolving dataset; culling the dynamically evolving dataset based on a threshold metric; and training an updated audio source separation model using the culled dynamically evolving dataset.

[0176] 46. ​​The method according to Examples 41-45, wherein the self-repeating training process further includes the steps of: calculating a first quality metric associated with a plurality of generated audio stems, the first quality metric providing a first performance measure of the trained audio source separation model; calculating a second quality metric associated with an improved audio stem, the second quality metric providing a performance measure of the updated audio source separation model; and comparing the second quality metric with the first quality metric and confirming that the second quality metric is superior to the first quality metric.

[0177] 47. The method according to Examples 41-46, wherein the trained audio source separation model comprises multiple datasets, each comprising labeled audio samples, and is trained using a training dataset, each of which is configured to train the audio source separation model to address a different source separation problem.

[0178] 48. The method according to Examples 41-47, wherein the multiple datasets include a speech training dataset comprising multiple labeled speech samples and / or a non-speech training dataset comprising multiple labeled music and / or noise data samples.

[0179] 49. The method according to Examples 41-48, wherein the self-repeating training process further includes the step of generating labeled audio samples from multiple audio stems generated for a self-repeating dataset.

[0180] 50. The method according to Examples 41-49, further comprising the step of generating multiple improved audio stems using a hierarchical branching sequence that includes the step of separating the source signal and the remaining complementary signals.

[0181] Where applicable, the various implementations provided by this disclosure may be implemented using hardware, software, or a combination of hardware and software. Where applicable, the various hardware and / or software components described herein may be combined into composite components comprising software, hardware, and / or both, without departing from the spirit of this disclosure. Where applicable, the various hardware and / or software components described herein may be separated into subcomponents comprising software, hardware, or both, without departing from the spirit of this disclosure.

[0182] The software described herein, including non-transient instructions, program code, and / or data, may be stored on one or more non-transient machine-readable media. It is also assumed that the software identified herein may be implemented using one or more general-purpose or dedicated computers and / or computer systems, which are networked and / or otherwise. Where applicable, the ordering of the various steps described herein may be modified, combined into compound steps, and / or separated into substeps to provide the features described herein. The implementations described above illustrate, but do not limit, the invention. It should also be understood that numerous modifications and variations are possible in accordance with the principles of the invention. Therefore, the scope of the invention is defined only by the following claims.

Claims

[Claim 1] The invention described herein.