Audio Source Separation for Noisy Single-Track Recordings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio source separation techniques are not optimized to generate high-quality audio stems from low-quality, single-track, noisy sound mixtures, particularly from older recordings.
Innovation Solution
A modified recurrent neural network (RNN) class model with specific manipulations, including removing redundant layers, omitting masking steps, applying windowing functions, and using a separation strength parameter, is trained with a self-iterative dataset generation loop and hierarchical mix bus schema to enhance audio source separation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If existing audio source separation techniques are used, then the process can separate audio components, but the output quality is insufficient for high-fidelity production
Solution Approach 1:
The patent applies parameter changes by modifying the RNN architecture (removing redundant layers, changing activation functions), adjusting hyperparameters (learning rates, batch sizes), and transforming input data (normalization, augmentation) to improve separation accuracy and output quality
Solution Approach 2:
The patent implements feedback mechanisms through self-iterative dataset generation where initial separation results are used to create training data, quality metrics evaluate output stems, and results feed back into model refinement to continuously improve separation accuracy
2Object-affected harmful factors
If traditional separation models are used, then processing can be performed, but artifacts and noise are not adequately reduced
Solution Approach 1:
The patent converts harmful artifacts and noise into beneficial training data by using self-iterative generation where initial separation results (even with artifacts) are processed to create synthetic clean audio for training, transforming the problem into a learning opportunity
Solution Approach 2:
The patent creates synthetic copies of audio data through self-iterative dataset generation, where processed audio is reused to train new models, effectively copying and refining separation results across multiple iterations to reduce artifacts
3Manufacturing precision
If complex processing pipelines are used, then separation quality may improve, but system complexity increases
Solution Approach 1:
The patent extracts and removes redundant layers from the RNN architecture, identifies and eliminates unnecessary masking steps, and simplifies the processing pipeline while maintaining separation quality through more efficient model design
Data Source
AI summary
Systems and methods for audio source separation include receiving an audio input stream including a mixture of audio signals generated from a plurality of audio sources; processing, through a trained audio source separation model, the audio input stream to generate a plurality of audio stems corresponding to one or more of the plurality of audio sources; updating, using a self-iterative processing and training system, the audio source separation model based at least in part on the plurality of audio stems; and re-processing, using the updated trained audio source separation model, the audio input stream to generate a plurality of enhanced audio stems.


