Dynamic Facial Feature Speaker Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speaker verification systems face challenges in robustness due to noise and ease of deception, relying solely on acoustic characteristics, which can lead to user frustration and security risks.
Innovation Solution
The system incorporates dynamic facial features by prompting users to utter a unique text challenge, recording both audio and video, and performing a comparison to a database of observable behaviors, enhancing security through the alignment of audio and video features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If speaker verification relies solely on acoustic characteristics, then the system is simple to implement, but robustness degrades in noisy environments and against voice playback attacks
Solution Approach 1:
The patent combines acoustic characteristics analysis with visual facial feature analysis into a unified speaker verification system. The system processes both audio signals and video feeds simultaneously, merging multiple biometric modalities to improve reliability and robustness against spoofing attacks while maintaining manageable system complexity through integrated processing architecture
Solution Approach 2:
The verification system uses a composite approach by combining multiple types of biometric data (acoustic voice characteristics and visual facial dynamics) into a composite verification model. This composite methodology enhances robustness by requiring multiple concurrent validations, making the system resistant to single-point failures or targeted attacks on individual biometric modalities
2Reliability
If the system uses only acoustic characteristics for verification, then the implementation is straightforward, but the system is vulnerable to voice playback attacks
Solution Approach 1:
The system merges acoustic verification with visual verification processes. While the acoustic component detects voice characteristics, the visual component simultaneously analyzes facial muscle movements and expressions to detect genuine speech production, creating a multi-layered security approach that thwarts voice playback attacks without excessive complexity
Solution Approach 2:
The visual facial analysis acts as an intermediary verification layer between the acoustic verification and the final authentication decision. By analyzing facial dynamics as an intermediate step, the system can detect spoofing attempts and make more secure verification decisions without directly complicating the core acoustic verification process
3Reliability
If traditional acoustic-only verification is used, then user interaction is simple, but user frustration increases when verification fails due to noise or connection issues
Solution Approach 1:
The system merges acoustic and visual verification channels to provide redundant verification paths. When acoustic verification fails or degrades due to noise or connection issues, the visual facial analysis channel can compensate and maintain verification accuracy, thereby reducing user frustration while keeping the interaction model relatively simple
Data Source
AI summary
Disclosed herein are systems, methods, and non-transitory computer-readable storage media for performing speaker verification. A system configured to practice the method receives a request to verify a speaker, generates a text challenge that is unique to the request, and, in response to the request, prompts the speaker to utter the text challenge. Then the system records a dynamic image feature of the speaker as the speaker utters the text challenge, and performs speaker verification based on the dynamic image feature and the text challenge. Recording the dynamic image feature of the speaker can include recording video of the speaker while speaking the text challenge. The dynamic feature can include a movement pattern of head, lips, mouth, eyes, and/or eyebrows of the speaker. The dynamic image feature can relate to phonetic content of the speaker speaking the challenge, speech prosody, and the speaker's facial expression responding to content of the challenge.


