Dynamic Challenge Utterance Generation for Speaker Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speaker verification systems are vulnerable to automated attacks, as they often rely on a limited set of recorded words or utterances, making it possible for unauthorized access using stolen speech.
Innovation Solution
Implementing a system that generates random utterances based on a large vocabulary and unique phonemes, phoneme clusters, and prosodic patterns to create challenge sentences that maximize speaker discrimination while minimizing length, and utilizing stealth enrollment through multi-platform automatic speech recognition engines to build user profiles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a limited set of recorded words or utterances is used in speaker verification systems, then the system is easier to implement and operate, but the system becomes vulnerable to automated attacks using stolen speech
Solution Approach 1:
The system dynamically generates challenge utterances in real-time using a large vocabulary and grammatical rules, rather than relying on a fixed set of pre-recorded words. This dynamic generation ensures that each verification challenge is unique and unpredictable, preventing automated attacks while maintaining system operability through automated challenge creation and evaluation
Solution Approach 2:
The system changes the parameter of utterance variety by using a large vocabulary and generating diverse sentence structures through grammatical rules. This transforms the verification process from using a limited set of fixed utterances to generating an effectively unlimited set of unique challenge sentences, thereby securing the system against replay attacks
2Reliability
If a large vocabulary and many sentences are used for random utterance generation, then speaker verification security is improved, but the system complexity increases
Solution Approach 1:
The system segments the challenge generation process into distinct components: a large vocabulary database, grammatical rules for sentence construction, and a random selection mechanism. This segmentation allows the system to manage complexity by organizing the large vocabulary into structured categories and applying systematic grammatical rules, making the overall system more tractable despite the large scale
Solution Approach 2:
The system employs a universal grammatical framework that can generate diverse challenge sentences from a single vocabulary set. This multi-functional approach allows the same grammatical rules to produce various sentence structures and meanings, reducing system complexity by reusing core components rather than creating separate mechanisms for each utterance type
3Measurement precision
If challenge sentences are customized for each speaker based on voice characteristics, then speaker discrimination accuracy is improved, but the processing time and system complexity increase
Solution Approach 1:
The system performs preliminary analysis of speaker voice characteristics during the enrollment phase, storing extracted acoustic features and pronunciation patterns in a user profile. This preliminary action enables the system to quickly generate customized challenge sentences during verification by referencing pre-stored speaker-specific information, rather than performing complex analysis in real-time
Solution Approach 2:
The system replaces complex real-time voice characteristic analysis with a streamlined process that uses pre-stored speaker profiles and automated grammar-based generation. This substitution reduces processing time by eliminating the need for complex mechanical analysis during verification, while still achieving high discrimination accuracy through speaker-specific challenge generation
Data Source
AI summary
Disclosed herein are systems, methods, and non-transitory computer-readable storage media relating to speaker verification. In one aspect, a system receives a first user identity from a second user, and, based on the identity, accesses voice characteristics. The system randomly generates a challenge sentence according to a rule and/or grammar, based on the voice characteristics, and prompts the second user to speak the challenge sentence. The system verifies that the second user is the first user if the spoken challenge sentence matches the voice characteristics. In an enrollment aspect, the system constructs an enrollment phrase that covers a minimum threshold of unique speech sounds based on speaker-distinctive phonemes, phoneme clusters, and prosody. Then user utters the enrollment phrase and extracts voice characteristics for the user from the uttered enrollment phrase. The system generates a user profile, based on the voice characteristics, for generating random challenge sentences according to a grammar.


