Voice Biometrics Replay Attack Detection Using Microphone Array
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice biometrics systems face challenges in detecting replay attacks, where a malicious party records and plays back an enrolled user's voice to gain unauthorized access, as high-quality loudspeakers can produce sound waves that are indistinguishable from the original voice, making it difficult to identify spoofing attempts.
Innovation Solution
The method involves using multiple microphones to receive speech signals with different frequency components, determining the position of each frequency component's source, and comparing these positions to identify if they differ by more than a threshold, indicating a potential replay attack, by utilizing techniques like generalized cross-correlation with phase transform (GCC-PHAT) to estimate time delays and angles of arrival.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high-quality loudspeakers are used to play back recorded voice, then the sound quality becomes indistinguishable from original voice, but the system becomes vulnerable to replay attacks
Solution Approach 1:
The patent transitions from analyzing only audio frequency domain to incorporating spatial domain information by using multiple microphones to capture directional data. The system estimates angles of arrival for different frequency components and uses this spatial dimension to detect replay attacks, thereby maintaining voice recognition accuracy while improving security.
Solution Approach 2:
The patent divides the speech signal into different frequency components (e.g., low-frequency and high-frequency parts) and separately estimates the angle of arrival for each component. In a genuine speech signal, all frequency components should originate from the same direction, but in a replay attack, different frequency components may come from different directions due to speaker characteristics. This segmentation approach enables detection of replay attacks while preserving accurate voice recognition.
2Reliability
If multiple microphones are used to detect frequency component positions, then replay attack detection capability is improved, but device complexity increases
Solution Approach 1:
The patent applies local quality by having different microphones or microphone pairs specialize in detecting specific frequency components. Each microphone or processing channel is optimized to analyze particular frequency ranges, and the system combines these localized analyses to achieve comprehensive replay attack detection. This approach improves detection accuracy while managing system complexity through specialized local processing.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach effectively differentiates between original voice inputs and replayed signals, enhancing the security of voice biometrics systems by accurately identifying and preventing unauthorized access attempts.
Implementation Method 1
receiving a speech signal at at least a first microphone and a second microphone
Implementation Method 2
utilizing techniques like generalized cross-correlation with phase transform (GCC-PHAT) to estimate time delays and angles of arrival
Data Source
AI summary
In order to detect a replay attack on a voice biometrics system, a speech signal is received at at least a first microphone and a second microphone. The speech signal has components at first and second frequencies. The method of detection comprises: obtaining information about a position of a source of the first frequency component of the speech signal, relative to the first and second microphones; obtaining information about a position of a source of the second frequency component of the speech signal, relative to the first and second microphones; comparing the position of the source of the first frequency component and the position of the source of the second frequency component; and determining that the speech signal may result from a replay attack if the position of the source of the first frequency component differs from the position of the source of the second frequency component by more than a threshold amount.


