Replay Spoofing Detection in Automatic Speaker Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current replay detection methods for automatic speaker verification systems fail to generalize well to unseen replay configurations, such as different background noises and recording devices, due to their reliance on spectrum-related features and classification techniques.
Innovation Solution
A neural network machine learning model is trained using genuine and replay sample data to optimize the classification of input biometric data, where results for genuine samples are closer to a genuine center and results for replay samples are further away, allowing for effective detection of replay attacks without transforming raw audio data into a spectrum.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If spectrum-related features are extracted and classification techniques are used for replay detection, then replay detection capability is achieved, but generalization to unseen replay configurations deteriorates
Solution Approach 1:
The patent transforms the replay detection problem from spectrum-based feature extraction to time-domain raw audio processing. By changing the fundamental parameter domain from frequency spectrum to time-domain waveforms and using neural network-based loss functions (additive noise loss, multiplicative noise loss) instead of traditional classification metrics, the system achieves better generalization to unseen replay configurations while maintaining detection accuracy
Solution Approach 2:
The patent replaces traditional mechanical classification systems (SVM, GMM, CNN classifiers) with a neural network-based optimization approach. Instead of extracting features and feeding them to classifiers, the system uses neural networks to directly process raw audio and optimize loss functions that measure the difference between genuine and replay audio characteristics, substituting classification mechanics with continuous optimization
2Reliability
If traditional classification techniques are applied to spectrum features, then binary classification is achieved, but robustness to different background noises and devices deteriorates
Solution Approach 1:
The patent changes the parameter space from spectrum features to time-domain raw audio and replaces classification metrics with noise-robust loss functions. The additive noise loss and multiplicative noise loss functions are specifically designed to be invariant to background noise and device characteristics, allowing the system to maintain reliability across different environmental conditions without requiring retraining or adaptation
Data Source
AI summary
Described herein are a system and techniques for detecting whether biometric data provided in an access request is genuine or a replay. In some embodiments, the system uses an machine learning model trained using genuine and replay sample data which is optimized in order to produce a result set in which results for the genuine samples are pulled closer to a genuine center and results for the replay samples are pushed away from the genuine center. Subjecting input biometric data (e.g., an audio sample) to the trained model results in a classification of the input biometric data as genuine or replay, which can then be used to determine whether or not to verify the input biometric data.


