Audio Source Separation for Noisy Single-Track Recordings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio source separation techniques are not optimized to generate high-quality audio stems from low-quality, single-track, noisy sound mixtures, particularly from older recordings.

Innovation Solution

A modified recurrent neural network (RNN) class model with specific manipulations, including removing redundant layers, omitting masking steps, applying windowing functions, and using a separation strength parameter, is trained with a self-iterative dataset generation loop and hierarchical mix bus schema to enhance audio source separation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If existing audio source separation techniques are used, then the process can separate audio components, but the output quality is insufficient for high-fidelity production

Engineering Contradiction:
Improveaudio stem qualityVSAvoidseparation accuracy
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent applies parameter changes by modifying the RNN architecture (removing redundant layers, changing activation functions), adjusting hyperparameters (learning rates, batch sizes), and transforming input data (normalization, augmentation) to improve separation accuracy and output quality

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements feedback mechanisms through self-iterative dataset generation where initial separation results are used to create training data, quality metrics evaluate output stems, and results feed back into model refinement to continuously improve separation accuracy

Inventive Principle:
Principle #23Feedback

2Object-affected harmful factors

If traditional separation models are used, then processing can be performed, but artifacts and noise are not adequately reduced

Engineering Contradiction:
Improveartifacts and noiseVSAvoidaudio fidelity
Core Design Contradiction:
Object-affected harmful factorsVSManufacturing precision

Solution Approach 1:

The patent converts harmful artifacts and noise into beneficial training data by using self-iterative generation where initial separation results (even with artifacts) are processed to create synthetic clean audio for training, transforming the problem into a learning opportunity

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

The patent creates synthetic copies of audio data through self-iterative dataset generation, where processed audio is reused to train new models, effectively copying and refining separation results across multiple iterations to reduce artifacts

Inventive Principle:
Principle #26Copying

3Manufacturing precision

If complex processing pipelines are used, then separation quality may improve, but system complexity increases

Engineering Contradiction:
Improveaudio stem qualityVSAvoidprocessing system complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent extracts and removes redundant layers from the RNN architecture, identifies and eliminates unnecessary masking steps, and simplifies the processing pipeline while maintaining separation quality through more efficient model design

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12431159B2Audio source separation systems and methods
Publication Date: 2025.09.30 WINGNUT FILMS PROD LTD
  • US12431159B2 patent drawing
  • US12431159B2 patent drawing
  • US12431159B2 patent drawing

AI summary

Systems and methods for audio source separation include receiving an audio input stream including a mixture of audio signals generated from a plurality of audio sources; processing, through a trained audio source separation model, the audio input stream to generate a plurality of audio stems corresponding to one or more of the plurality of audio sources; updating, using a self-iterative processing and training system, the audio source separation model based at least in part on the plurality of audio stems; and re-processing, using the updated trained audio source separation model, the audio input stream to generate a plurality of enhanced audio stems.