Speaker Recognition Using Multiple Speech Samples

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speaker recognition methods in telephone calls are error-prone and lack sufficient reliability, often confusing unknown speakers with known ones due to insufficient accuracy in matching speech samples.

Innovation Solution

A method that involves obtaining and classifying multiple speech samples from unknown speakers to create speaker-dependent classes, extracting and combining speaker information, and comparing it with stored target speaker information using Gaussian Mixture Models and feature vectors to determine identity based on similarity thresholds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional speaker recognition methods using single speech samples and Gaussian Mixture Models are used, then the processing speed is maintained, but the reliability and accuracy of speaker identification deteriorates due to error-prone matching

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidmethod complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the speaker identification process into distinct phases: obtaining multiple speech samples, classifying them into speaker-dependent classes, extracting speaker information from each class, combining the extracted information, and finally comparing with stored profiles. This segmentation allows each step to be optimized independently, improving overall reliability while managing complexity through structured processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary classification of speech samples into speaker-dependent classes before the actual speaker identification comparison. This preliminary action organizes the speech data in advance, creating structured speaker profiles that enhance the reliability of subsequent identification operations without adding complexity to the core comparison algorithm.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If multiple speech samples are obtained and classified into speaker-dependent classes, then the accuracy of speaker recognition is improved, but the processing time and computational complexity increases

Engineering Contradiction:
Improvespeaker matching accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent divides the processing of multiple speech samples into separate speaker-dependent classes, where each class is processed independently for speaker information extraction. This segmentation enables parallel processing of different speech samples, improving measurement precision through multiple data points while reducing total processing time through concurrent operations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts and combines speaker information from multiple speech samples, using more data than the conventional single-sample approach. This excessive action of processing multiple samples provides redundant information that improves accuracy, while the systematic combination method ensures that the additional processing does not linearly increase time consumption.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If speaker information is extracted and combined from multiple speaker-dependent classes, then the reliability of speaker recognition is enhanced, but the computational resources required increase

Engineering Contradiction:
Improvespeaker recognition reliabilityVSAvoidcomputational energy
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent merges speaker information extracted from multiple speaker-dependent classes into a unified speaker profile through combination operations. This merging consolidates redundant computational work, where the combined information from multiple classes provides enhanced reliability without requiring proportionally increased computational energy, as shared processing steps are performed once and reused.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS9043207B2Speaker recognition from telephone calls
Publication Date: 2015.05.26 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9043207B2 patent drawing
  • US9043207B2 patent drawing

AI summary

The present invention relates to a method for speaker recognition, comprising the steps of obtaining and storing speaker information for at least one target speaker; obtaining a plurality of speech samples from a plurality of telephone calls from at least one unknown speaker; classifying the speech samples according to the at least one unknown speaker thereby providing speaker-dependent classes of speech samples; extracting speaker information for the speech samples of each of the speaker-dependent classes of speech samples; combining the extracted speaker information for each of the speaker-dependent classes of speech samples; comparing the combined extracted speaker information for each of the speaker-dependent classes of speech samples with the stored speaker information for the at least one target speaker to obtain at least one comparison result; and determining whether one of the at least one unknown speakers is identical with the at least one target speaker based on the at least one comparison result.