Multi-Speaker Speech Recognition via Voiceprint Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional single-channel speech recognition methods fail to accurately recognize multiple speakers and their overlapping voices in multi-speaker and multi-channel scenarios, such as conferences or customer service settings, as they are designed for single-channel and single-speaker environments.
Innovation Solution
A multi-channel and multi-speaker speech recognition method that performs speech activity detection, separates overlapping speech segments, extracts voiceprint features, and clusters audio segments based on these features to identify individual speakers, enabling accurate recognition of multiple speakers and their corresponding speech.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional single-channel speech recognition method is used, then the system is simple and easy to operate, but it cannot recognize respective voices of multiple speakers especially the overlapping parts
Solution Approach 1:
The patent applies segmentation by dividing the mixed speech signal into separate speech segments through speech separation. The system segments overlapping speech from multiple speakers into individual audio streams, then processes each segment separately through speech recognition and voiceprint extraction, enabling accurate identification of each speaker's contribution in multi-speaker scenarios
Solution Approach 2:
The patent introduces voiceprint feature vectors as an intermediary element to bridge the gap between separated speech segments and speaker identification. The voiceprint extraction module generates unique acoustic fingerprints for each speaker, which serve as mediators to match separated speech segments with their corresponding speakers, enabling accurate speaker attribution without direct complex modeling
2Measurement precision
If speech separation and clustering operations are performed, then recognition accuracy for multiple speakers is improved, but processing time and computational complexity increase
Solution Approach 1:
The patent applies preliminary action by performing speech activity detection and speech separation before speech recognition. By pre-separating overlapping speech into distinct segments and pre-extracting voiceprint features, the system prepares the data in advance, which reduces the computational burden during the actual recognition phase and enables more accurate processing of each speaker's speech independently
Solution Approach 2:
The patent segments the processing pipeline into distinct stages: speech activity detection, speech separation, voiceprint extraction, clustering, and speech recognition. This segmentation allows each module to specialize in a specific task, improving overall accuracy while enabling parallel processing of different speech segments, which helps mitigate the time penalty through efficient resource utilization
3Loss of information
If speech separation is performed on speech segments with multiple speakers, then overlapping voices are separated, but the device complexity and processing steps increase
Solution Approach 1:
The patent segments the mixed speech signal into separate speaker-specific audio segments through speech separation. This segmentation preserves speaker identity by isolating each speaker's voice characteristics in separate time-frequency regions, then processes each segment independently through voiceprint extraction and clustering to maintain accurate speaker attribution throughout the pipeline
Solution Approach 2:
The patent transforms the speech separation problem from a temporal domain challenge into a joint time-frequency domain solution. By analyzing speech in the frequency domain and separating speakers based on their spectral characteristics and temporal patterns simultaneously, the system preserves speaker identity more effectively while managing complexity through dimensionality transformation rather than单纯 temporal processing
Data Source
AI summary
A speech recognition method. The method includes: performing speech activity detection on speech data to obtain multiple speech segments; determining, for each of the speech segments, a number of speakers involved in the each of the speech segments; for each of at least one of the speech segments with the determined number greater than 1: performing speech separation on the each of at least one of the speech segments to obtain multiple audio segments; performing speech recognition on each of the audio segments to obtain respective first speech recognition results for the audio segments; performing feature extraction on each of the audio segments to obtain respective voiceprint feature vectors; and performing clustering on the audio segments with respect to the speakers to obtain a clustering result; and obtaining a second speech recognition result for the speech data based on the clustering result and the respective first speech recognition results.


