Spoofprint Embedding for Robust Voice Biometric Spoof Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice biometric systems are vulnerable to spoofing attacks from advanced speech synthesis technologies, failing to detect unknown spoofing techniques and lacking generalization ability.
Innovation Solution
A neural network architecture is employed for spoof detection, utilizing embedding extractors to differentiate between voiceprint and spoofprint features, with a large margin cosine loss function to maximize variance between genuine and spoofed classes, and minimize intra-class variance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional voice biometric systems are used for speaker recognition, then voice matching can be performed, but the systems become vulnerable to spoofing attacks from advanced speech synthesis technologies
Solution Approach 1:
The patent segments the feature extraction process into two distinct pathways: voiceprint extraction for speaker identification and spoofprint extraction for spoof detection. This segmentation allows the system to independently optimize features for each purpose, with spoofprint features specifically designed to detect spoofing artifacts while voiceprint features focus on speaker characteristics. The neural network architecture includes separate embedding extractors for voiceprints and spoofprints, enabling simultaneous operation of both functions without interference.
Solution Approach 2:
The patent introduces a new dimensional space for spoof detection by extracting spoofprint features in addition to traditional voiceprint features. This adds a new dimension to the authentication process, transforming it from a single-dimensional speaker verification task into a multi-dimensional system that simultaneously performs speaker identification and spoof detection. The spoofprint embedding space is specifically designed to capture spoofing artifacts that are orthogonal to speaker identity features.
2Measurement precision
If speech synthesis tools are used to generate synthesized speech, then voice features can be mimicked, but spoofing artifacts are introduced that can be detected
Solution Approach 1:
The patent extracts spoofing artifacts as a separate, distinct feature set from the audio signal through dedicated spoofprint embedding extractors. Instead of attempting to detect spoofs by analyzing modifications to existing voiceprint features, the system extracts spoofprint features that specifically represent spoofing characteristics. This extraction approach isolates the detection task from the complexity of speech synthesis variations, allowing the neural network to focus on identifying spoofing artifacts without being overwhelmed by the diversity of synthesis techniques.
Solution Approach 2:
The neural network architecture is designed with universal components that serve multiple functions. The same neural network backbone processes both voiceprint and spoofprint extraction, sharing computational resources and parameters. The embedding extractors are trained jointly on both tasks, allowing the system to perform speaker verification and spoof detection simultaneously with a single model, rather than requiring separate specialized systems for each function.
Data Source
AI summary
Embodiments described herein provide for systems and methods for implementing a neural network architecture for spoof detection in audio signals. The neural network architecture contains a layers defining embedding extractors that extract embeddings from input audio signals. Spoofprint embeddings are generated for particular system enrollees to detect attempts to spoof the enrollee's voice. Optionally, voiceprint embeddings are generated for the system enrollees to recognize the enrollee's voice. The voiceprints are extracted using features related to the enrollee's voice. The spoofprints are extracted using features related to features of how the enrollee speaks and other artifacts. The spoofprints facilitate detection of efforts to fool voice biometrics using synthesized speech (e.g., deepfakes) that spoof and emulate the enrollee's voice.


