Audio Activity Detection Using Time-Encoding for Always-On Speech
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-controlled devices face challenges in power consumption and delay in processing audio signals due to the need for continuous operation of speech recognition modules, leading to potential loss of audio data when transitioning from standby mode to active processing.
Innovation Solution
An activity detector using a first time-encoding modulator generating a PWM signal and a second time-encoding modulator generating a clock signal, with a time-decoding converter and activity monitor to determine signal activity, allowing for low-power operation and rapid detection of voice activity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If the speech recognition module is kept always-on to enable voice control functionality, then the device can respond to voice commands at any time, but power consumption increases significantly
Solution Approach 1:
The audio processing system is segmented into two distinct modules: a low-power voice activity detection (VAD) module that continuously monitors audio signals, and a high-power speech recognition module that processes audio only when VAD detects speech activity. This segmentation allows the device to maintain always-on voice control capability while minimizing power consumption by keeping the speech recognition module inactive during non-speech periods.
2Use of energy by moving object
If the speech recognition module is turned off to save power, then power consumption decreases, but there is a delay when turning it back on causing audio data loss
Solution Approach 1:
The voice activity detection module performs preliminary monitoring of audio signals continuously in a low-power state. When speech activity is detected, the VAD module triggers the speech recognition module to activate. This preliminary detection mechanism ensures that the speech recognition module is activated at the precise moment speech occurs, eliminating activation delay and preventing audio data loss while maintaining power-saving operation during non-speech periods.
3Loss of information
If the back-end processing is activated immediately upon detecting audio activity, then no audio data is lost, but power consumption increases during standby periods
Solution Approach 1:
The voice activity detection module serves as an intermediary between the continuous audio input and the speech recognition module. The VAD module continuously monitors audio signals in a low-power state and intelligently determines when speech activity occurs. Based on VAD's detection, the speech recognition module is selectively activated only when necessary. This intermediary mechanism ensures complete audio data capture during speech events while minimizing power consumption during non-speech standby periods.
Data Source
AI summary
This application relates an activity detector (100) for detecting signal activity in an input audio signal (SIN), such as may be used for always-on speech detection. The activity detector has a first time-encoding modulator (TEM) 101 including a first hysteretic comparator (201) for generating a PWM (pulse-width modulation) signal based on the input audio signal. A second TEM (103) having a second hysteretic comparator (401) is arranged to receive a reference voltage (VMID) and generate a clock signal (SCLK). A time-decoding converter (102) receives the clock signal and generates count values of a number of cycles of the clock signal in periods defined by the PWM signal. An activity monitor (104) is responsive to a count signal (SCT) from the TDC 102 to determine whether the input audio signal comprises signal activity above a defined threshold.


