Discrete Token Audio Source Separation via Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio systems struggle to accurately isolate and enhance individual sound sources from complex audio mixtures, often introducing artifacts or distortions that impact perceptual quality and downstream tasks like speech recognition.
Innovation Solution
The use of machine learning and discrete tokens to estimate different sound sources from audio mixtures, where audio input is converted into discrete tokens and a trained machine learning model identifies sound sources corresponding to subsets of these tokens.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional audio systems use traditional signal processing methods to separate sound sources, then the system complexity remains low, but the separation accuracy deteriorates and artifacts are introduced
Solution Approach 1:
The patent replaces traditional mechanical signal processing methods with a machine learning-based system. A neural network model processes audio mixtures to separate sound sources, substituting conventional filtering and spectral analysis with learned representations that achieve superior separation accuracy without introducing artifacts.
Solution Approach 2:
The system transforms the audio separation problem by changing the parameter space from traditional time-frequency domain processing to a learned latent space representation. The neural network learns optimal parameter transformations that enable accurate sound source separation while maintaining computational feasibility.
2Reliability
If conventional methods process audio mixtures directly in time domain, then computational complexity remains low, but the ability to isolate individual sources deteriorates in challenging acoustic environments
Solution Approach 1:
The patent moves the processing from the traditional time domain to a transformed domain using neural network features. By projecting audio data into a higher-dimensional feature space learned by the model, the system achieves robust separation in challenging acoustic environments while managing computational complexity through efficient architecture design.
3Measurement precision
If neural networks are trained on artificial mixtures of isolated speech, then training data generation is simple, but the accuracy of separating real-world sound sources deteriorates
Solution Approach 1:
The system performs preliminary processing by converting audio mixtures into discrete token representations before feeding them to the separation model. This pre-tokenization step creates a more manageable and informative input representation that improves accuracy while maintaining training data generation simplicity through automated transcription and tokenization pipelines.
Data Source
AI summary
A system and method are disclosed. Audio input comprising the mixed audio signals is received by one or more client devices. The audio input is converted into a plurality of discrete tokens. A plurality of sound sources, each corresponding to a subset of discrete tokens of a plurality of subsets of discrete tokens, is determined using a trained machine learning model.


