Speaker Authentication Using Discriminative Section Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speaker recognition algorithms are sensitive to variations in utterance methods and vocal organ structures, limiting their ability to accurately discriminate between speakers.
Innovation Solution
A method and apparatus for speaker authentication that dynamically matches input speaker features to pre-enrolled features by selecting and weighting discriminable sections, aligning phonemes, and dropping non-discriminative sections, using neural networks to extract and estimate speaker features and discriminable sections.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a scoring scheme is used for speaker recognition, then speaker discrimination is attempted, but the scheme becomes sensitive to utterance method variations of the same speaker, reducing discrimination accuracy
Solution Approach 1:
The patent segments the speaker verification process into multiple scoring schemes (first scoring scheme for utterance method insensitivity, second scoring scheme for vocal organ sensitivity). Each scheme processes different aspects of speaker characteristics independently, then their results are combined. This segmentation allows the system to separately handle utterance variations and vocal organ differences, resolving the contradiction between sensitivity to utterance variations and speaker discrimination accuracy.
Solution Approach 2:
The patent applies different quality characteristics to different parts of the scoring process. The first scoring scheme is designed with local quality optimized for insensitivity to utterance method changes, while the second scoring scheme has local quality optimized for sensitivity to vocal organ differences. By assigning different local qualities to different scoring components, the system achieves both utterance insensitivity and vocal organ sensitivity simultaneously.
2Loss of information
If all speaker features are used for matching, then comprehensive speaker information is captured, but features with low discriminability (e.g., short pauses) reduce authentication accuracy
Solution Approach 1:
The patent extracts and removes non-discriminative features (such as short pauses with low discriminable speaker sections) from the speaker feature set before authentication. By taking out these harmful features that reduce authentication accuracy while preserving comprehensive speaker information through the multi-scoring approach, the system resolves the contradiction between information completeness and authentication precision.
Solution Approach 2:
The patent discards features with low discriminability (short pauses) from the authentication process while recovering and maintaining comprehensive speaker information through the combination of multiple scoring schemes. The discarded features are compensated for by the robustness of the multi-scoring approach, which maintains authentication accuracy without relying on non-discriminative features.
Data Source
AI summary
A speaker authentication method and apparatus may extract input speaker features corresponding to a plurality of frames of an input speech of an object, estimate discriminable speaker sections corresponding to the plurality of frames, and dynamically match the input speaker features to pre-enrolled enrolled speaker features based on the discriminable speaker section.


