Online Audio Source Separation Using Pre-computed Reference Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech denoising techniques perform poorly with non-stationary background noise, as they cannot effectively model noise signals with rapidly changing spectral profiles, leading to suboptimal separation of speech from background noise in teleconferencing and audio/video chatting.

Innovation Solution

The implementation of online source separation techniques using pre-computed reference data and algorithms like Probabilistic Latent Component Analysis (PLCA) for real-time separation of audio signals, allowing for the modeling of time-varying spectral profiles and effective separation of speech from non-stationary noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional speech denoising techniques using a single spectral profile are used, then the system is simple to implement, but it performs poorly with non-stationary noise

Engineering Contradiction:
Improvespeech denoising performanceVSAvoidnoise modeling complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the noise signal into multiple spectral profiles organized in a hierarchical structure (tree), where each node represents a cluster of similar noise frames. This segmentation allows the system to adapt to non-stationary noise by selecting appropriate spectral profiles from different segments, resolving the contradiction between maintaining simplicity and improving denoising performance.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic noise modeling by continuously updating the spectral profile tree structure based on incoming noise frames. The system dynamically adapts to changing noise characteristics by adding, removing, or merging spectral profiles in real-time, enabling reliable performance with non-stationary noise while managing complexity through efficient data structures.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If a single spectral profile is used to model noise, then the computational complexity is low, but the ability to handle rapidly changing noise spectra is poor

Engineering Contradiction:
Improvenoise spectrum adaptabilityVSAvoidspectral modeling complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal noise model that can handle multiple types of noise conditions through a single hierarchical spectral profile structure. The tree-based organization allows the same data structure to represent both stationary and non-stationary noise, providing versatility across different noise scenarios without requiring separate models for each noise type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent transitions from a single-dimensional spectral profile representation to a multi-dimensional hierarchical structure. By organizing spectral profiles in a tree with multiple levels and clusters, the system adds dimensional complexity that enables better representation of non-stationary noise while managing computational load through the hierarchical organization.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If real-time online separation is performed, then the speech output is timely and relevant, but the processing requirements increase

Engineering Contradiction:
Improvereal-time processing speedVSAvoidcomputational power
Core Design Contradiction:
ProductivityVSPower

Solution Approach 1:

The patent performs preliminary organization of noise spectral profiles into a hierarchical tree structure during idle periods or initialization phases. This pre-processing allows the real-time separation algorithm to efficiently query and update the existing structure without building it from scratch, reducing computational power requirements during time-critical real-time operation while maintaining high productivity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9966088B2Online source separation
Publication Date: 2018.05.08 ADOBE INC
  • US9966088B2 patent drawing
  • US9966088B2 patent drawing
  • US9966088B2 patent drawing

AI summary

Online source separation may include receiving a sound mixture that includes first audio data from a first source and second audio data from a second source. Online source separation may further include receiving pre-computed reference data corresponding to the first source. Online source separation may also include performing online separation of the second audio data from the first audio data based on the pre-computed reference data.