Speaker Recognition Using Multiple Speech Samples
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speaker recognition methods in telephone calls are error-prone and lack sufficient reliability, often confusing unknown speakers with known ones due to insufficient accuracy in matching speech samples.
Innovation Solution
A method that involves obtaining and classifying multiple speech samples from unknown speakers to create speaker-dependent classes, extracting and combining speaker information, and comparing it with stored target speaker information using Gaussian Mixture Models and feature vectors to determine identity based on similarity thresholds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional speaker recognition methods using single speech samples and Gaussian Mixture Models are used, then the processing speed is maintained, but the reliability and accuracy of speaker identification deteriorates due to error-prone matching
Solution Approach 1:
The patent segments the speaker identification process into distinct phases: obtaining multiple speech samples, classifying them into speaker-dependent classes, extracting speaker information from each class, combining the extracted information, and finally comparing with stored profiles. This segmentation allows each step to be optimized independently, improving overall reliability while managing complexity through structured processing.
Solution Approach 2:
The patent performs preliminary classification of speech samples into speaker-dependent classes before the actual speaker identification comparison. This preliminary action organizes the speech data in advance, creating structured speaker profiles that enhance the reliability of subsequent identification operations without adding complexity to the core comparison algorithm.
2Measurement precision
If multiple speech samples are obtained and classified into speaker-dependent classes, then the accuracy of speaker recognition is improved, but the processing time and computational complexity increases
Solution Approach 1:
The patent divides the processing of multiple speech samples into separate speaker-dependent classes, where each class is processed independently for speaker information extraction. This segmentation enables parallel processing of different speech samples, improving measurement precision through multiple data points while reducing total processing time through concurrent operations.
Solution Approach 2:
The patent extracts and combines speaker information from multiple speech samples, using more data than the conventional single-sample approach. This excessive action of processing multiple samples provides redundant information that improves accuracy, while the systematic combination method ensures that the additional processing does not linearly increase time consumption.
3Reliability
If speaker information is extracted and combined from multiple speaker-dependent classes, then the reliability of speaker recognition is enhanced, but the computational resources required increase
Solution Approach 1:
The patent merges speaker information extracted from multiple speaker-dependent classes into a unified speaker profile through combination operations. This merging consolidates redundant computational work, where the combined information from multiple classes provides enhanced reliability without requiring proportionally increased computational energy, as shared processing steps are performed once and reused.
Data Source
AI summary
The present invention relates to a method for speaker recognition, comprising the steps of obtaining and storing speaker information for at least one target speaker; obtaining a plurality of speech samples from a plurality of telephone calls from at least one unknown speaker; classifying the speech samples according to the at least one unknown speaker thereby providing speaker-dependent classes of speech samples; extracting speaker information for the speech samples of each of the speaker-dependent classes of speech samples; combining the extracted speaker information for each of the speaker-dependent classes of speech samples; comparing the combined extracted speaker information for each of the speaker-dependent classes of speech samples with the stored speaker information for the at least one target speaker to obtain at least one comparison result; and determining whether one of the at least one unknown speakers is identical with the at least one target speaker based on the at least one comparison result.

