Joint Speaker Diarization and Separation Using Time-Invariant Embeddings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional sound processing systems face challenges in accurately performing speaker separation and diarization, especially in multi-talker conversational speech, due to the complexity of designing and training multiple neural networks for sequential or concurrent tasks, which increases computational demands and reduces efficiency in real-time processing.

Innovation Solution

A deep neural network is trained jointly for both speaker diarization and separation tasks, using a speaker-independent layer and a speaker-biased layer to process audio mixtures, eliminating the need for additional neural networks and reducing computational requirements by extracting time-frequency activity regions of each speaker.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple neural networks are used for sequential or concurrent speaker separation and diarization tasks, then task accuracy is improved, but device complexity and computational demands increase

Engineering Contradiction:
Improvespeaker separation and diarization accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges speaker separation and diarization tasks into a single neural network architecture. The network simultaneously performs both functions by processing audio mixtures through shared convolutional layers followed by task-specific output layers, eliminating the need for separate networks for each task.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The neural network is designed as a universal system that can perform multiple functions - both speaker separation and diarization - using a single unified architecture. The network takes audio mixtures as input and produces both separated speaker signals and diarization information through its multi-output structure.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If multiple neural networks are used for sequential or concurrent speaker separation and diarization tasks, then task accuracy is improved, but computational requirements and processing efficiency deteriorate

Engineering Contradiction:
Improvespeaker separation and diarization accuracyVSAvoidprocessing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

By combining both tasks into one neural network, the patent reduces redundant computational operations. The shared convolutional layers process the audio mixture once, and both separation and diarization outputs are generated from this common representation, improving computational efficiency.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The universal network architecture processes audio mixtures in a single pass to produce both separated speaker signals and diarization information, eliminating the need for sequential processing steps and multiple separate network executions, thereby improving overall processing efficiency.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If traditional source separation systems are used to isolate specific sound types, then separation capability is improved, but adaptability to changing target sounds at test time deteriorates

Engineering Contradiction:
Improvetarget sound isolation capabilityVSAvoidflexibility in target sound selection
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The network incorporates dynamic conditioning mechanisms that allow the target speaker to be changed during operation. By using conditioning vectors that can be updated at test time, the system adapts its separation behavior to isolate different speakers dynamically without requiring retraining.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes its separation parameters by modifying conditioning vectors rather than retraining the entire network. This allows flexible adaptation to different target speakers by adjusting the conditioning input, enabling the system to switch between isolating different speakers based on changing requirements.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240304205A1System and Method for Audio Processing using Time-Invariant Speaker Embeddings
Publication Date: 2024.09.12 MITSUBISHI ELECTRIC RESEARCH LABORATORIES INC
  • US20240304205A1 patent drawing
  • US20240304205A1 patent drawing
  • US20240304205A1 patent drawing

AI summary

A system and method for sound processing for performing multi-talker conversation analysis is provided. The sound processing system includes a deep neural network trained for processing audio segments of an audio mixture of the multi-talker conversation. The deep neural network includes a speaker-independent layer that produces a speaker-independent output, and a speaker-biased layer applied once independently to each of the audio segments for each multiple speakers of the audio mixture. The deep neural network also processes a time-invariant embedding by individually assigning each application of the speaker-biased layer to a corresponding speaker by inputting the corresponding time-invariant speaker embedding. The deep neural network thus produces data indicative of time-frequency activity regions of each speaker of the multiple speakers in the audio mixture from a combination of speaker-biased outputs.