De-mixing Filter Selection for Multi-Speaker Audio Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Blind Source Separation (BSS) systems face delays and high computational loads due to the need for frequent re-training when multiple speakers move within a space, as the statistical characteristics of their audio change rapidly, making it difficult to separate their voices effectively without significant processing resources.

Innovation Solution

Initial training with a generalized voice for multiple locations generates sets of de-mixing filters for different positions, which are stored and selected based on the speaker's location, allowing for simultaneous multiple user speech recognition without additional training, eliminating delays and reducing computational load.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If Blind Source Separation (BSS) systems perform frequent re-training when speakers move, then separation quality is maintained, but computational load and processing delays increase significantly

Engineering Contradiction:
Improveseparation qualityVSAvoidprocessing delays
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system pre-calculates and stores de-mixing filters for multiple predetermined speaker positions during an initialization phase. When speakers move, the system selects from these pre-computed filters based on detected speaker positions, eliminating the need for frequent re-training and thus reducing processing delays while maintaining separation quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The continuous spatial environment is segmented into discrete predetermined positions, each with its own pre-computed de-mixing filter. This segmentation allows the system to handle speaker movement by selecting from discrete filter sets rather than continuously re-computing filters, reducing computational load and processing time.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If Blind Source Separation (BSS) systems perform frequent re-training when speakers move, then separation quality is maintained, but computational resources are excessively consumed

Engineering Contradiction:
Improveseparation qualityVSAvoidcomputational load
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The system pre-calculates and stores de-mixing filters for multiple predetermined speaker positions during an initialization phase. When speakers move, the system selects from these pre-computed filters based on detected speaker positions, eliminating the need for frequent re-training and thus reducing processing delays while maintaining separation quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of continuously computing new de-mixing filters when speakers move, the system creates copies of pre-computed filter sets for different predetermined positions. The appropriate filter copy is selected based on current speaker positions, avoiding the computational expense of re-training while maintaining separation effectiveness.

Inventive Principle:
Principle #26Copying

3Object-affected harmful factors

If noise cancellation microphones are used to isolate user voice, then background noise is reduced, but complex audio processing is required

Engineering Contradiction:
Improvebackground noiseVSAvoidaudio processing complexity
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

The system uses multiple microphones not only for noise cancellation but also for determining speaker positions through time-delay analysis. This multi-functional approach allows the same hardware to serve both noise reduction and source localization purposes, reducing overall system complexity while maintaining effective noise isolation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9817634B2Distinguishing speech from multiple users in a computer interaction
Publication Date: 2017.11.14 INTEL CORP
  • US9817634B2 patent drawing
  • US9817634B2 patent drawing
  • US9817634B2 patent drawing

AI summary

Speech from multiple users is distinguished. In one example, an apparatus has a sensor to determine a position of a speaker, a microphone array to receive audio from the speaker and from other simultaneous audio sources, and a processor to select a pre-determined filter based on the determined position and to apply the selected filter to the received audio to separate the audio from the speaker from the audio from the other simultaneous audio sources.