Discrete Token Audio Source Separation via Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio systems struggle to accurately isolate and enhance individual sound sources from complex audio mixtures, often introducing artifacts or distortions that impact perceptual quality and downstream tasks like speech recognition.

Innovation Solution

The use of machine learning and discrete tokens to estimate different sound sources from audio mixtures, where audio input is converted into discrete tokens and a trained machine learning model identifies sound sources corresponding to subsets of these tokens.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional audio systems use traditional signal processing methods to separate sound sources, then the system complexity remains low, but the separation accuracy deteriorates and artifacts are introduced

Engineering Contradiction:
Improvesound source separation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical signal processing methods with a machine learning-based system. A neural network model processes audio mixtures to separate sound sources, substituting conventional filtering and spectral analysis with learned representations that achieve superior separation accuracy without introducing artifacts.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system transforms the audio separation problem by changing the parameter space from traditional time-frequency domain processing to a learned latent space representation. The neural network learns optimal parameter transformations that enable accurate sound source separation while maintaining computational feasibility.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If conventional methods process audio mixtures directly in time domain, then computational complexity remains low, but the ability to isolate individual sources deteriorates in challenging acoustic environments

Engineering Contradiction:
Improverobustness in challenging acoustic environmentsVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent moves the processing from the traditional time domain to a transformed domain using neural network features. By projecting audio data into a higher-dimensional feature space learned by the model, the system achieves robust separation in challenging acoustic environments while managing computational complexity through efficient architecture design.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If neural networks are trained on artificial mixtures of isolated speech, then training data generation is simple, but the accuracy of separating real-world sound sources deteriorates

Engineering Contradiction:
Improvesound source estimation accuracyVSAvoidtraining data generation ease
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The system performs preliminary processing by converting audio mixtures into discrete token representations before feeding them to the separation model. This pre-tokenization step creates a more manageable and informative input representation that improves accuracy while maintaining training data generation simplicity through automated transcription and tokenization pipelines.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250054500A1Using machine learning and discrete tokens to estimate different sound sources from audio mixtures
Publication Date: 2025.02.13 GOOGLE LLC
  • US20250054500A1 patent drawing
  • US20250054500A1 patent drawing
  • US20250054500A1 patent drawing

AI summary

A system and method are disclosed. Audio input comprising the mixed audio signals is received by one or more client devices. The audio input is converted into a plurality of discrete tokens. A plurality of sound sources, each corresponding to a subset of discrete tokens of a plurality of subsets of discrete tokens, is determined using a trained machine learning model.