Integrated Signal Processing for Blind and Target Speaker Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current blind sound source separation and target speaker extraction technologies are constructed as independent systems, making it difficult to use them appropriately in situations where prior information is unknown, and they suffer from permutation issues between utterances.
Innovation Solution
A signal processing device and method that integrates blind sound source separation and target speaker extraction using a neural network with a conversion unit, weighting unit, and mask estimation unit, allowing for switching between blind and target speaker extraction based on the presence of auxiliary information, and updates parameters through multi-task learning.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If blind sound source separation is used, then speaker separation is possible without prior information, but permutation problem occurs between utterances
Solution Approach 1:
The system performs preliminary speaker embedding extraction and stores speaker characteristics before processing mixed speech. By pre-computing speaker embeddings and maintaining speaker identity representations, the system can reliably track speakers across utterances without permutation issues, while still operating in blind separation mode without requiring prior speaker information.
2Reliability
If target speaker extraction is used, then permutation problem between utterances is solved, but it cannot be applied when speakers are not known in advance
Solution Approach 1:
The system creates a universal framework that can perform both blind sound source separation and target speaker extraction through a single integrated architecture. By incorporating speaker embedding extraction and similarity computation capabilities, the system can adapt to either mode of operation depending on whether auxiliary speaker information is provided, making it universally applicable to both known and unknown speaker scenarios.
3Measurement precision
If blind sound source separation and target speaker extraction are constructed as independent systems, then each can be optimized for its specific purpose, but they cannot be appropriately used with one model
Solution Approach 1:
The system merges blind sound source separation and target speaker extraction into a single integrated model. The architecture combines speaker embedding extraction, mixed speech encoding, and separation functions in one unified neural network. This allows the system to maintain optimized performance for both tasks while reducing complexity by eliminating the need for separate independent systems.
Data Source
AI summary
A signal processing device according to an embodiment of the present invention includes: a conversion unit configured to convert an input mixed acoustic signal into a plurality of first internal states, a weighting unit configured to generate a second internal state which is a weighted sum of the plurality of first internal states based on auxiliary information regarding an acoustic signal of a target sound source when the auxiliary information is input, and generate the second internal state by selecting one of the plurality of first internal states when the auxiliary information is not input, and a mask estimation unit configured to estimate a mask based on the second internal state.


