Voice Print Categorization via Neural Network Compression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional biometric authentication systems using voice prints face challenges with high storage requirements and latency due to large voice print files, leading to potential fraud and inaccurate authentication, especially with users who are skilled at self-identifying or imitating voice prints.

Innovation Solution

The system employs an AI model, such as a neural network, to categorize voice prints into distinct categories (e.g., 'sheep,' 'goats,' and 'wolves') based on their self-identification and impersonation likelihood, allowing for efficient compression and storage, with dynamic threshold adjustments for authentication, thereby reducing storage needs and improving authentication speed and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If voice prints are stored in high capacity storage systems, then authentication accuracy is maintained, but storage requirements increase and latency increases

Engineering Contradiction:
Improveauthentication accuracyVSAvoidstorage capacity
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The voice print data is segmented into two parts: (1) a compressed representation stored in high-capacity storage systems, and (2) a categorization label stored separately. The compression mechanism divides the voice print into essential features that can be efficiently stored, reducing the storage footprint while maintaining authentication accuracy through the categorization information that guides the authentication process.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes the parameter of voice print representation by applying compression algorithms that transform the raw voice print data into a condensed format. The compression mechanism modifies the data parameters to reduce storage requirements while preserving the essential authentication characteristics through the categorization labels that encode information about voice print properties.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If voice prints are compressed to reduce storage, then storage capacity and latency are improved, but data loss occurs affecting authentication quality

Engineering Contradiction:
Improvestorage capacityVSAvoiddata quality
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The compression mechanism extracts only the essential features and characteristics from the voice print data for storage. By identifying and storing only the critical information needed for authentication (such as unique voice patterns and categorical properties), the system reduces data size while maintaining sufficient quality for accurate authentication, discarding redundant information that does not contribute to authentication accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system creates a compressed copy of the voice print that contains the essential authentication information. This compressed representation is a strategic copy that preserves the most important characteristics through the categorization labels, allowing authentication to proceed accurately without requiring the full original data to be stored.

Inventive Principle:
Principle #26Copying

3Quantity of substance

If compression is applied to voice prints, then storage requirements are reduced, but authentication time increases due to processing needs

Engineering Contradiction:
Improvestorage capacityVSAvoidauthentication latency
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The compression and categorization of voice prints is performed as a preliminary action during the enrollment phase. By pre-compressing and categorizing voice prints before they are needed for authentication, the system eliminates the need for compression during the authentication process itself, thereby reducing authentication latency while still achieving storage efficiency.

Inventive Principle:
Principle #10Preliminary action

4Quantity of substance

If standard compression is used for voice prints, then storage capacity is improved, but fraud risk increases due to data loss

Engineering Contradiction:
Improvestorage capacityVSAvoidfraud risk
Core Design Contradiction:
Quantity of substanceVSObject-affected harmful factors

Solution Approach 1:

The categorization labels provide feedback information about the compressed voice print data, enabling the authentication system to adjust its evaluation criteria based on the category. This feedback mechanism allows the system to compensate for any data loss during compression by using the categorical information to guide the authentication decision, thereby maintaining security and reducing fraud risk.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The authentication threshold and evaluation criteria are made dynamic based on the voice print category. Different categories have different authentication requirements and thresholds, allowing the system to adapt to the specific characteristics of each voice print type. This dynamic approach compensates for compression-induced data loss by adjusting the authentication criteria to match the compressed data's characteristics, thereby maintaining security.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11756555B2Biometric authentication through voice print categorization using artificial intelligence
Publication Date: 2023.09.12 NICE LTD
  • US11756555B2 patent drawing
  • US11756555B2 patent drawing
  • US11756555B2 patent drawing

AI summary

A system is provided to categorize voice prints during a voice authentication. The system includes a processor and a computer readable medium operably coupled thereto, to perform voice authentication operations which include receiving an enrollment of a user in the biometric authentication system, requesting a first voice print comprising a sample of a voice of the user, receiving the first voice print of the user during the enrollment, accessing a plurality of categorizations of the voice prints for the voice authentication, wherein each of the plurality of categorizations comprises a portion of the voice prints based on a plurality of similarity scores of distinct voice prints in the portion to a plurality of other voice prints, determining, using a hidden layer of a neural network, one of the plurality of categorizations for the first voice print, and encoding the first voice print with the one of the plurality of categorizations.