Audio Activity Detection Using Time-Encoding Modulators
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-controlled devices face challenges in power consumption and delay in processing audio signals due to the continuous operation of speech recognition modules, leading to potential loss of audio data when transitioning from a low-power standby mode to active processing.
Innovation Solution
An activity detector using a first time-encoding modulator generating a PWM signal and a second time-encoding modulator generating a clock signal, with a time-decoding converter and activity monitor to determine signal activity, allowing the system to switch between low-power and high-resolution modes efficiently.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If the speech recognition module is kept continuously active to enable always-on voice control, then the device can respond to voice commands at any time, but power consumption increases significantly
Solution Approach 1:
The speech processing system is divided into two distinct modules: a voice activity detection (VAD) module that operates continuously at low power to monitor for speech, and a speech recognition module that operates only when speech is detected. This segmentation allows the system to maintain always-on voice control capability while significantly reducing overall power consumption by keeping the high-power recognition module dormant during non-speech periods.
2Use of energy by moving object
If the speech recognition module is switched off to save power, then power consumption is reduced, but there is a delay when turning it back on causing audio data loss
Solution Approach 1:
The voice activity detection module performs preliminary monitoring of the audio signal continuously, detecting speech activity before the speech recognition module needs to process it. This preliminary detection allows the system to wake up the speech recognition module in advance, ensuring it is fully operational and ready to process audio data without delay or loss when speech is actually present.
3Reliability
If an ADC is operated continually to convert analogue audio signals for digital processing, then audio data is not lost during wake-up, but power consumption increases
Solution Approach 1:
The signal processing chain is segmented into separate functional blocks with different power requirements. The VAD module operates in a low-power mode using minimal processing, while the ADC and speech recognition module operate at full power only when needed. This segmentation allows the system to maintain audio data integrity during speech events without the penalty of continuous high-power operation.
Solution Approach 2:
Instead of continuous operation, the ADC and speech recognition module are activated periodically based on speech detection events. The VAD module monitors continuously and triggers periodic activation of the higher-power components only when speech is detected, thereby maintaining data integrity during speech while minimizing overall power consumption during non-speech periods.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This solution reduces power consumption and minimizes data loss by enabling the speech processing module only when signal activity is detected, ensuring timely and efficient processing of audio signals while maintaining low power operation.
Implementation Method 1
a first time-encoding modulator comprising a first hysteretic comparator for generating a PWM (pulse-width modulation) signal based on the input audio signal
Implementation Method 2
a time-decoding converter configured to receive the clock signal, generate count values of a number of cycles of the clock signal in periods defined by the PWM signal
Data Source
AI summary
This application relates an activity detector (100) for detecting signal activity in an input audio signal (SIN), such as may be used for always-on speech detection. The activity detector has a first time-encoding modulator (TEM) 101 including a first hysteretic comparator (201) for generating a PWM (pulse-width modulation) signal based on the input audio signal. A second TEM (103) having a second hysteretic comparator (401) is arranged to receive a reference voltage (VMID) and generate a clock signal (SCLK). A time-decoding converter (102) receives the clock signal and generates count values of a number of cycles of the clock signal in periods defined by the PWM signal. An activity monitor (104) is responsive to a count signal (SCT) from the TDC 102 to determine whether the input audio signal comprises signal activity above a defined threshold.


