Speaker Recognition Using Voice Feature Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice recognition techniques for authenticating speakers require a large number of voice samples in a database, leading to increased calculation volume and reduced accuracy when the utterance time is short.

Innovation Solution

A method that groups voice information of unspecified speakers similar in feature to a registered speaker's voice information, using these groups to calculate similarity with a subject voice signal, thereby reducing the number of voice samples needed for authentication without increasing calculation volume.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the number of voice information samples in the database is increased to improve authentication accuracy, then the accuracy in identifying the authentic speaker is improved, but the calculation volume increases

Engineering Contradiction:
Improveauthentication accuracyVSAvoidcalculation volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent segments the large database of voice information into multiple clusters, where each cluster contains voice samples grouped by similarity in voice features. During authentication, only the cluster corresponding to the claimed speaker is retrieved and used for comparison, rather than processing the entire database. This segmentation reduces the calculation volume while maintaining authentication accuracy by focusing computations on relevant data subsets.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary clustering of voice information during the database setup phase, organizing samples into groups based on voice feature similarity before authentication occurs. This preliminary action creates a structured database where speaker-specific clusters are pre-formed, enabling efficient retrieval during authentication without requiring full database scans, thus reducing real-time calculation volume while preserving accuracy.

Inventive Principle:
Principle #10Preliminary action

2Quantity of substance

If the number of voice information samples is reduced to decrease calculation volume, then the calculation volume is reduced, but the accuracy in identifying the authentic speaker deteriorates

Engineering Contradiction:
Improvecalculation volumeVSAvoidauthentication accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent applies local quality by creating speaker-specific clusters where voice samples are grouped according to their similarity to each speaker's voice features. Each cluster contains only the relevant samples needed for that speaker's authentication, giving different parts of the database different compositions tailored to specific speakers. This allows the system to use a smaller, focused subset of data for each authentication decision, reducing overall calculation volume while maintaining accuracy through targeted sample selection.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If voice information of unspecified speakers is grouped by similarity to registered speakers, then the accuracy in identifying authentic speakers is improved without increasing calculation volume, but the device complexity increases

Engineering Contradiction:
Improveauthentication accuracyVSAvoiddatabase structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates simplified representations (clusters) of speaker voice profiles by grouping similar voice samples. Each cluster serves as a condensed copy of a speaker's voice characteristics, containing only the essential similar samples needed for authentication. This copying approach maintains authentication accuracy by preserving key voice features while reducing the overall data structure complexity compared to storing and processing individual unrelated samples.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11315573B2Speaker recognizing method, speaker recognizing apparatus, recording medium recording speaker recognizing program, database making method, database making apparatus, and recording medium recording database making program
Publication Date: 2022.04.26 PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
  • US11315573B2 patent drawing
  • US11315573B2 patent drawing
  • US11315573B2 patent drawing

AI summary

A speaker recognizing method includes acquiring subject identification information that is identification information of an authentic person who the subject speaker claims to be, calculating a first feature value representing a feature value of the subject voice signal, selecting a group including pieces of the voice information associated with the subject identification information, from the first database, calculating degrees of similarity between the pieces of the voice information included in the selected group and the first feature value and a subject degree of similarity representing a degree of similarity between the voice information associated with the subject identification information, the voice information being stored in the second database, and the first feature value, calculating a rank of the subject degree of similarity in the calculated degrees of similarity, and when the rank is smaller than a given first rank, determining the subject speaker to be the authentic person.