Multi-Speaker Speech Recognition via Voiceprint Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional single-channel speech recognition methods fail to accurately recognize multiple speakers and their overlapping voices in multi-speaker and multi-channel scenarios, such as conferences or customer service settings, as they are designed for single-channel and single-speaker environments.

Innovation Solution

A multi-channel and multi-speaker speech recognition method that performs speech activity detection, separates overlapping speech segments, extracts voiceprint features, and clusters audio segments based on these features to identify individual speakers, enabling accurate recognition of multiple speakers and their corresponding speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional single-channel speech recognition method is used, then the system is simple and easy to operate, but it cannot recognize respective voices of multiple speakers especially the overlapping parts

Engineering Contradiction:
Improvecapability to recognize multiple speakersVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the mixed speech signal into separate speech segments through speech separation. The system segments overlapping speech from multiple speakers into individual audio streams, then processes each segment separately through speech recognition and voiceprint extraction, enabling accurate identification of each speaker's contribution in multi-speaker scenarios

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces voiceprint feature vectors as an intermediary element to bridge the gap between separated speech segments and speaker identification. The voiceprint extraction module generates unique acoustic fingerprints for each speaker, which serve as mediators to match separated speech segments with their corresponding speakers, enabling accurate speaker attribution without direct complex modeling

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If speech separation and clustering operations are performed, then recognition accuracy for multiple speakers is improved, but processing time and computational complexity increase

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing speech activity detection and speech separation before speech recognition. By pre-separating overlapping speech into distinct segments and pre-extracting voiceprint features, the system prepares the data in advance, which reduces the computational burden during the actual recognition phase and enables more accurate processing of each speaker's speech independently

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the processing pipeline into distinct stages: speech activity detection, speech separation, voiceprint extraction, clustering, and speech recognition. This segmentation allows each module to specialize in a specific task, improving overall accuracy while enabling parallel processing of different speech segments, which helps mitigate the time penalty through efficient resource utilization

Inventive Principle:
Principle #1Segmentation

3Loss of information

If speech separation is performed on speech segments with multiple speakers, then overlapping voices are separated, but the device complexity and processing steps increase

Engineering Contradiction:
Improvepreservation of speaker identityVSAvoidnumber of processing steps
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the mixed speech signal into separate speaker-specific audio segments through speech separation. This segmentation preserves speaker identity by isolating each speaker's voice characteristics in separate time-frequency regions, then processes each segment independently through voiceprint extraction and clustering to maintain accurate speaker attribution throughout the pipeline

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the speech separation problem from a temporal domain challenge into a joint time-frequency domain solution. By analyzing speech in the frequency domain and separating speakers based on their spectral characteristics and temporal patterns simultaneously, the system preserves speaker identity more effectively while managing complexity through dimensionality transformation rather than单纯 temporal processing

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12183362B1Speech recognition
Publication Date: 2024.12.31 MASHANG CONSUMER FINANCE CO LTD
  • US12183362B1 patent drawing
  • US12183362B1 patent drawing
  • US12183362B1 patent drawing

AI summary

A speech recognition method. The method includes: performing speech activity detection on speech data to obtain multiple speech segments; determining, for each of the speech segments, a number of speakers involved in the each of the speech segments; for each of at least one of the speech segments with the determined number greater than 1: performing speech separation on the each of at least one of the speech segments to obtain multiple audio segments; performing speech recognition on each of the audio segments to obtain respective first speech recognition results for the audio segments; performing feature extraction on each of the audio segments to obtain respective voiceprint feature vectors; and performing clustering on the audio segments with respect to the speakers to obtain a clustering result; and obtaining a second speech recognition result for the speech data based on the clustering result and the respective first speech recognition results.