Integrated Signal Processing for Blind and Target Speaker Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current blind sound source separation and target speaker extraction technologies are constructed as independent systems, making it difficult to use them appropriately in situations where prior information is unknown, and they suffer from permutation issues between utterances.

Innovation Solution

A signal processing device and method that integrates blind sound source separation and target speaker extraction using a neural network with a conversion unit, weighting unit, and mask estimation unit, allowing for switching between blind and target speaker extraction based on the presence of auxiliary information, and updates parameters through multi-task learning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If blind sound source separation is used, then speaker separation is possible without prior information, but permutation problem occurs between utterances

Engineering Contradiction:
Improveability to separate speakers without prior informationVSAvoidpermutation consistency between utterances
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system performs preliminary speaker embedding extraction and stores speaker characteristics before processing mixed speech. By pre-computing speaker embeddings and maintaining speaker identity representations, the system can reliably track speakers across utterances without permutation issues, while still operating in blind separation mode without requiring prior speaker information.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If target speaker extraction is used, then permutation problem between utterances is solved, but it cannot be applied when speakers are not known in advance

Engineering Contradiction:
Improvepermutation consistency between utterancesVSAvoidability to handle unknown speakers
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system creates a universal framework that can perform both blind sound source separation and target speaker extraction through a single integrated architecture. By incorporating speaker embedding extraction and similarity computation capabilities, the system can adapt to either mode of operation depending on whether auxiliary speaker information is provided, making it universally applicable to both known and unknown speaker scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If blind sound source separation and target speaker extraction are constructed as independent systems, then each can be optimized for its specific purpose, but they cannot be appropriately used with one model

Engineering Contradiction:
Improveseparation performance for specific purposeVSAvoidnumber of independent systems
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system merges blind sound source separation and target speaker extraction into a single integrated model. The architecture combines speaker embedding extraction, mixed speech encoding, and separation functions in one unified neural network. This allows the system to maintain optimized performance for both tasks while reducing complexity by eliminating the need for separate independent systems.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11978471B2Signal processing apparatus, learning apparatus, signal processing method, learning method and program
Publication Date: 2024.05.07 NIPPON TELEGRAPH & TELEPHONE CORP
  • US11978471B2 patent drawing
  • US11978471B2 patent drawing
  • US11978471B2 patent drawing

AI summary

A signal processing device according to an embodiment of the present invention includes: a conversion unit configured to convert an input mixed acoustic signal into a plurality of first internal states, a weighting unit configured to generate a second internal state which is a weighted sum of the plurality of first internal states based on auxiliary information regarding an acoustic signal of a target sound source when the auxiliary information is input, and generate the second internal state by selecting one of the plurality of first internal states when the auxiliary information is not input, and a mask estimation unit configured to estimate a mask based on the second internal state.