Speaker Diarization via Blind Source Separation and Spatial Matrices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speaker diarization technologies face challenges in accurately distinguishing speakers with similar voices using single microphone systems and those at close angles with multi-microphone systems, especially in reverberant environments, leading to low diarization accuracy due to component aging and the need for prior knowledge of microphone array arrangements.
Innovation Solution
An audio signal processing method that employs a multi-microphone system, utilizing blind source separation and spatial characteristic matrices to improve speaker diarization accuracy by clustering speakers without requiring prior knowledge of microphone array arrangements, and incorporating preset audio features to handle scenarios with speakers at close angles and those who move.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If single microphone-based speaker diarization is used, then device complexity is reduced, but speaker distinction accuracy deteriorates when speakers have similar voices
Solution Approach 1:
The patent combines multiple microphones into an array system to capture spatial information from different positions. By merging the capabilities of multiple microphones and integrating spatial characteristic matrices with audio features, the system achieves superior speaker distinction accuracy compared to single microphone systems.
2Measurement precision
If multi-microphone-based speaker diarization system is used, then speaker distinction capability is improved, but system complexity increases and component aging affects performance
Solution Approach 1:
The system performs self-calibration by automatically estimating the arrangement of microphone array components through spatial characteristic matrix calculations. This self-service capability eliminates the need for manual calibration and compensates for component aging effects, maintaining accurate speaker distinction without increasing operational complexity.
3Measurement precision
If prior knowledge of microphone array arrangement is required, then spatial processing accuracy is improved, but adaptability deteriorates when components age or reposition
Solution Approach 1:
The system dynamically adapts to changes in microphone array configuration by continuously estimating spatial characteristics from incoming audio signals. Rather than relying on fixed prior knowledge, the system updates its understanding of microphone positions in real-time, maintaining accuracy even when components age or reposition.
4Speed
If conventional speaker diarization is used in reverberant environments, then processing speed is maintained, but diarization accuracy deteriorates due to reverb effects
Solution Approach 1:
The patent introduces spatial dimensionality to the speaker diarization process by incorporating spatial characteristic matrices that capture directional information from the microphone array. This additional spatial dimension enables the system to distinguish speakers based on their positions in the acoustic space, effectively separating direct speech from reverberant reflections and improving accuracy in reverberant environments.
Data Source
Figure 1A
Figure 1B
Figure 2A
AI summary
An audio signal processing method and a product are provided. The method includes: receiving N channels of observed signals collected by a microphone array, and performing blind source separation on the N channels of observed signals to obtain M channels of source signals and M demixing matrices, where the M channels of source signals are in a one-to-one correspondence with the M demixing matrices, N is an integer greater than or equal to 2, and M is an integer greater than or equal to 1 (S101); obtaining a spatial characteristic matrix corresponding to the N channels of observed signals, where the spatial characteristic matrix is used to represent a correlation between the N channels of observed signals (S102); obtaining a preset audio feature of each of the M channels of source signals (S103); and determining, based on the preset audio feature of each channel of source signal, the M demixing matrices, and the spatial characteristic matrix, a speaker quantity and a speaker identity corresponding to the N channels of observed signals (S104). The method helps improve speaker diarization accuracy.