Audio Activity Detection Using Time-Encoding for Always-On Speech

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice-controlled devices face challenges in power consumption and delay in processing audio signals due to the need for continuous operation of speech recognition modules, leading to potential loss of audio data when transitioning from standby mode to active processing.

Innovation Solution

An activity detector using a first time-encoding modulator generating a PWM signal and a second time-encoding modulator generating a clock signal, with a time-decoding converter and activity monitor to determine signal activity, allowing for low-power operation and rapid detection of voice activity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If the speech recognition module is kept always-on to enable voice control functionality, then the device can respond to voice commands at any time, but power consumption increases significantly

Engineering Contradiction:
Improvevoice control availabilityVSAvoidpower consumption
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The audio processing system is segmented into two distinct modules: a low-power voice activity detection (VAD) module that continuously monitors audio signals, and a high-power speech recognition module that processes audio only when VAD detects speech activity. This segmentation allows the device to maintain always-on voice control capability while minimizing power consumption by keeping the speech recognition module inactive during non-speech periods.

Inventive Principle:
Principle #1Segmentation

2Use of energy by moving object

If the speech recognition module is turned off to save power, then power consumption decreases, but there is a delay when turning it back on causing audio data loss

Engineering Contradiction:
Improvepower consumptionVSAvoidactivation delay
Core Design Contradiction:
Use of energy by moving objectVSLoss of time

Solution Approach 1:

The voice activity detection module performs preliminary monitoring of audio signals continuously in a low-power state. When speech activity is detected, the VAD module triggers the speech recognition module to activate. This preliminary detection mechanism ensures that the speech recognition module is activated at the precise moment speech occurs, eliminating activation delay and preventing audio data loss while maintaining power-saving operation during non-speech periods.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If the back-end processing is activated immediately upon detecting audio activity, then no audio data is lost, but power consumption increases during standby periods

Engineering Contradiction:
Improveaudio data completenessVSAvoidpower consumption
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The voice activity detection module serves as an intermediary between the continuous audio input and the speech recognition module. The VAD module continuously monitors audio signals in a low-power state and intelligently determines when speech activity occurs. Based on VAD's detection, the speech recognition module is selectively activated only when necessary. This intermediary mechanism ensures complete audio data capture during speech events while minimizing power consumption during non-speech standby periods.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11558706B2Activity detection
Publication Date: 2023.01.17 CIRRUS LOGIC INC
  • US11558706B2 patent drawing
  • US11558706B2 patent drawing
  • US11558706B2 patent drawing

AI summary

This application relates an activity detector (100) for detecting signal activity in an input audio signal (SIN), such as may be used for always-on speech detection. The activity detector has a first time-encoding modulator (TEM) 101 including a first hysteretic comparator (201) for generating a PWM (pulse-width modulation) signal based on the input audio signal. A second TEM (103) having a second hysteretic comparator (401) is arranged to receive a reference voltage (VMID) and generate a clock signal (SCLK). A time-decoding converter (102) receives the clock signal and generates count values of a number of cycles of the clock signal in periods defined by the PWM signal. An activity monitor (104) is responsive to a count signal (SCT) from the TDC 102 to determine whether the input audio signal comprises signal activity above a defined threshold.