Speaker Recognition Using Voice Feature Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition techniques for authenticating speakers require a large number of voice samples in a database, leading to increased calculation volume and reduced accuracy when the utterance time is short.
Innovation Solution
A method that groups voice information of unspecified speakers similar in feature to a registered speaker's voice information, using these groups to calculate similarity with a subject voice signal, thereby reducing the number of voice samples needed for authentication without increasing calculation volume.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the number of voice information samples in the database is increased to improve authentication accuracy, then the accuracy in identifying the authentic speaker is improved, but the calculation volume increases
Solution Approach 1:
The patent segments the large database of voice information into multiple clusters, where each cluster contains voice samples grouped by similarity in voice features. During authentication, only the cluster corresponding to the claimed speaker is retrieved and used for comparison, rather than processing the entire database. This segmentation reduces the calculation volume while maintaining authentication accuracy by focusing computations on relevant data subsets.
Solution Approach 2:
The patent performs preliminary clustering of voice information during the database setup phase, organizing samples into groups based on voice feature similarity before authentication occurs. This preliminary action creates a structured database where speaker-specific clusters are pre-formed, enabling efficient retrieval during authentication without requiring full database scans, thus reducing real-time calculation volume while preserving accuracy.
2Quantity of substance
If the number of voice information samples is reduced to decrease calculation volume, then the calculation volume is reduced, but the accuracy in identifying the authentic speaker deteriorates
Solution Approach 1:
The patent applies local quality by creating speaker-specific clusters where voice samples are grouped according to their similarity to each speaker's voice features. Each cluster contains only the relevant samples needed for that speaker's authentication, giving different parts of the database different compositions tailored to specific speakers. This allows the system to use a smaller, focused subset of data for each authentication decision, reducing overall calculation volume while maintaining accuracy through targeted sample selection.
3Measurement precision
If voice information of unspecified speakers is grouped by similarity to registered speakers, then the accuracy in identifying authentic speakers is improved without increasing calculation volume, but the device complexity increases
Solution Approach 1:
The patent creates simplified representations (clusters) of speaker voice profiles by grouping similar voice samples. Each cluster serves as a condensed copy of a speaker's voice characteristics, containing only the essential similar samples needed for authentication. This copying approach maintains authentication accuracy by preserving key voice features while reducing the overall data structure complexity compared to storing and processing individual unrelated samples.
Data Source
AI summary
A speaker recognizing method includes acquiring subject identification information that is identification information of an authentic person who the subject speaker claims to be, calculating a first feature value representing a feature value of the subject voice signal, selecting a group including pieces of the voice information associated with the subject identification information, from the first database, calculating degrees of similarity between the pieces of the voice information included in the selected group and the first feature value and a subject degree of similarity representing a degree of similarity between the voice information associated with the subject identification information, the voice information being stored in the second database, and the first feature value, calculating a rank of the subject degree of similarity in the calculated degrees of similarity, and when the rank is smaller than a given first rank, determining the subject speaker to be the authentic person.


