De-mixing Filter Selection for Multi-Speaker Audio Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Blind Source Separation (BSS) systems face delays and high computational loads due to the need for frequent re-training when multiple speakers move within a space, as the statistical characteristics of their audio change rapidly, making it difficult to separate their voices effectively without significant processing resources.
Innovation Solution
Initial training with a generalized voice for multiple locations generates sets of de-mixing filters for different positions, which are stored and selected based on the speaker's location, allowing for simultaneous multiple user speech recognition without additional training, eliminating delays and reducing computational load.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If Blind Source Separation (BSS) systems perform frequent re-training when speakers move, then separation quality is maintained, but computational load and processing delays increase significantly
Solution Approach 1:
The system pre-calculates and stores de-mixing filters for multiple predetermined speaker positions during an initialization phase. When speakers move, the system selects from these pre-computed filters based on detected speaker positions, eliminating the need for frequent re-training and thus reducing processing delays while maintaining separation quality.
Solution Approach 2:
The continuous spatial environment is segmented into discrete predetermined positions, each with its own pre-computed de-mixing filter. This segmentation allows the system to handle speaker movement by selecting from discrete filter sets rather than continuously re-computing filters, reducing computational load and processing time.
2Measurement precision
If Blind Source Separation (BSS) systems perform frequent re-training when speakers move, then separation quality is maintained, but computational resources are excessively consumed
Solution Approach 1:
The system pre-calculates and stores de-mixing filters for multiple predetermined speaker positions during an initialization phase. When speakers move, the system selects from these pre-computed filters based on detected speaker positions, eliminating the need for frequent re-training and thus reducing processing delays while maintaining separation quality.
Solution Approach 2:
Instead of continuously computing new de-mixing filters when speakers move, the system creates copies of pre-computed filter sets for different predetermined positions. The appropriate filter copy is selected based on current speaker positions, avoiding the computational expense of re-training while maintaining separation effectiveness.
3Object-affected harmful factors
If noise cancellation microphones are used to isolate user voice, then background noise is reduced, but complex audio processing is required
Solution Approach 1:
The system uses multiple microphones not only for noise cancellation but also for determining speaker positions through time-delay analysis. This multi-functional approach allows the same hardware to serve both noise reduction and source localization purposes, reducing overall system complexity while maintaining effective noise isolation.
Data Source
AI summary
Speech from multiple users is distinguished. In one example, an apparatus has a sensor to determine a position of a speaker, a microphone array to receive audio from the speaker and from other simultaneous audio sources, and a processor to select a pre-determined filter based on the determined position and to apply the selected filter to the received audio to separate the audio from the speaker from the audio from the other simultaneous audio sources.


