Digital Microphone Sampling Rate Control for Voice Anti-Spoofing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio systems struggle to differentiate between live and recorded speech, particularly in the context of voice biometric systems, due to limitations in detecting ultrasound components, which can lead to increased power consumption and vulnerability to replay attacks.
Innovation Solution
A method is introduced where a digital microphone's sampling rate is dynamically adjusted: a lower rate for normal operation to conserve power and a higher rate for detecting ultrasound components when a trigger phrase is spoken, allowing for effective anti-spoofing detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the sampling rate of the analog-digital converter is continuously maintained at a high rate to detect ultrasound components for spoof detection, then the reliability of spoof detection is improved, but the power consumption increases
Solution Approach 1:
The sampling rate of the analog-digital converter is dynamically adjusted based on operational context. The system switches between a first sampling rate (suitable for spoof detection) and a second sampling rate (for normal operation), rather than maintaining a fixed high sampling rate continuously. This dynamic adjustment resolves the contradiction by enabling high-reliability spoof detection only when necessary, thereby reducing overall power consumption.
Solution Approach 2:
The system changes the sampling rate parameter of the analog-digital converter based on detected speech presence. When speech is detected, the sampling rate is set to a first value optimized for ultrasound component analysis; when no speech is present, it switches to a second, lower sampling rate. This parameter change enables the system to maintain spoof detection capability while minimizing power consumption during idle periods.
2Use of energy by moving object
If the sampling rate is reduced to lower power consumption, then the power consumption decreases, but the ability to detect ultrasound components for spoof detection deteriorates
Solution Approach 1:
The sampling rate is made dynamic rather than static, allowing the system to adapt between power-saving mode (lower sampling rate) and detection mode (higher sampling rate). This resolves the contradiction by ensuring that the reduced sampling rate does not permanently degrade spoof detection capability, as the system can switch to the higher rate when detection is needed.
Solution Approach 2:
The system performs preliminary speech detection using the lower sampling rate, and only when speech is detected does it switch to the higher sampling rate for actual spoof detection analysis. This preliminary action allows the system to maintain low power consumption while ensuring that ultrasound components are properly captured when needed for security verification.
3Use of energy by moving object
If the sampling rate is dynamically adjusted between first and second rates, then the power consumption is optimized, but the device complexity increases
Solution Approach 1:
The audio processing circuit is designed to perform multiple functions: it can operate in a power-saving mode with lower sampling rate and switch to a detection mode with higher sampling rate when speech is present. This multi-functionality is achieved within a single integrated circuit that handles both normal audio processing and spoof detection, avoiding the need for separate dedicated hardware for each function and thus limiting the increase in device complexity.
Data Source
AI summary
An audio system receives an audio signal from a digital microphone, which has an analog-digital converter with a controllable sampling rate. In response to a determination that a predetermined trigger phrase is not detected in the decimated audio signal, the sampling rate of the analog-digital converter in the digital microphone is controlled such that the audio signal has a first sample rate. In response to a determination that the predetermined trigger phrase is detected in the decimated signal, the sampling rate of the analog-digital converter in the digital microphone is controlled such that the audio signal has a second sample rate higher than the first sample rate, and the audio signal is applied to a spoof detection circuit, to determine whether the received signal contains live speech or replayed speech.


