Dynamic Challenge Utterance Generation for Speaker Verification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speaker verification systems are vulnerable to automated attacks, as they often rely on a limited set of recorded words or utterances, making it possible for unauthorized access using stolen speech.

Innovation Solution

Implementing a system that generates random utterances based on a large vocabulary and unique phonemes, phoneme clusters, and prosodic patterns to create challenge sentences that maximize speaker discrimination while minimizing length, and utilizing stealth enrollment through multi-platform automatic speech recognition engines to build user profiles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a limited set of recorded words or utterances is used in speaker verification systems, then the system is easier to implement and operate, but the system becomes vulnerable to automated attacks using stolen speech

Engineering Contradiction:
Improveease of operationVSAvoidsecurity
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system dynamically generates challenge utterances in real-time using a large vocabulary and grammatical rules, rather than relying on a fixed set of pre-recorded words. This dynamic generation ensures that each verification challenge is unique and unpredictable, preventing automated attacks while maintaining system operability through automated challenge creation and evaluation

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the parameter of utterance variety by using a large vocabulary and generating diverse sentence structures through grammatical rules. This transforms the verification process from using a limited set of fixed utterances to generating an effectively unlimited set of unique challenge sentences, thereby securing the system against replay attacks

Inventive Principle:
Principle #35Parameter changes

2Reliability

If a large vocabulary and many sentences are used for random utterance generation, then speaker verification security is improved, but the system complexity increases

Engineering Contradiction:
ImprovesecurityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the challenge generation process into distinct components: a large vocabulary database, grammatical rules for sentence construction, and a random selection mechanism. This segmentation allows the system to manage complexity by organizing the large vocabulary into structured categories and applying systematic grammatical rules, making the overall system more tractable despite the large scale

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system employs a universal grammatical framework that can generate diverse challenge sentences from a single vocabulary set. This multi-functional approach allows the same grammatical rules to produce various sentence structures and meanings, reducing system complexity by reusing core components rather than creating separate mechanisms for each utterance type

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If challenge sentences are customized for each speaker based on voice characteristics, then speaker discrimination accuracy is improved, but the processing time and system complexity increase

Engineering Contradiction:
Improvespeaker discrimination accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of speaker voice characteristics during the enrollment phase, storing extracted acoustic features and pronunciation patterns in a user profile. This preliminary action enables the system to quickly generate customized challenge sentences during verification by referencing pre-stored speaker-specific information, rather than performing complex analysis in real-time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system replaces complex real-time voice characteristic analysis with a streamlined process that uses pre-stored speaker profiles and automated grammar-based generation. This substitution reduces processing time by eliminating the need for complex mechanical analysis during verification, while still achieving high discrimination accuracy through speaker-specific challenge generation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10121476B2System and method for generating challenge utterances for speaker verification
Publication Date: 2018.11.06 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10121476B2 patent drawing
  • US10121476B2 patent drawing
  • US10121476B2 patent drawing

AI summary

Disclosed herein are systems, methods, and non-transitory computer-readable storage media relating to speaker verification. In one aspect, a system receives a first user identity from a second user, and, based on the identity, accesses voice characteristics. The system randomly generates a challenge sentence according to a rule and/or grammar, based on the voice characteristics, and prompts the second user to speak the challenge sentence. The system verifies that the second user is the first user if the spoken challenge sentence matches the voice characteristics. In an enrollment aspect, the system constructs an enrollment phrase that covers a minimum threshold of unique speech sounds based on speaker-distinctive phonemes, phoneme clusters, and prosody. Then user utters the enrollment phrase and extracts voice characteristics for the user from the uttered enrollment phrase. The system generates a user profile, based on the voice characteristics, for generating random challenge sentences according to a grammar.