Cascade Audio Spotting System for Power-Constrained Keyword Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Audio controlled devices face challenges in reducing power consumption while maintaining accurate detection of spoken keywords or audio events, leading to high false acceptance and false rejection rates in noisy environments.
Innovation Solution
A cascade audio spotting system with multiple modules operating sequentially, where initial modules consume less power and have lower performance, and later modules consume more power and have higher performance, ensuring overall performance without unnecessary activation of high-power subsystems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a single high-performance audio processing module is used to detect audio events accurately, then detection accuracy is improved, but power consumption increases
Solution Approach 1:
The audio processing system is divided into multiple sequential modules with increasing performance capabilities. The cascade structure segments the detection task across modules of varying complexity, allowing the system to use only the necessary processing power for each situation rather than always deploying the full high-performance module.
Solution Approach 2:
The system applies partial action by using lower-power modules for routine detection and reserving high-power modules only when needed. This prevents excessive power consumption by activating full-performance processing only when the lower-power modules fail to detect or when detection confidence is insufficient.
2Use of energy by moving object
If a low-power audio processing module is used to reduce power consumption, then power consumption is reduced, but false acceptance and false rejection rates increase
Solution Approach 1:
The cascade structure introduces intermediate processing stages between low-power and high-power modules. When the low-power module detects an audio event, intermediate modules verify the detection before final confirmation, acting as mediators that reduce false acceptances and rejections without requiring constant high-power processing.
Solution Approach 2:
The system implements feedback mechanisms where detection results from lower-power modules are evaluated and fed back to determine whether higher-power modules should be activated. This feedback loop ensures that power consumption is optimized while maintaining reliability by triggering high-power processing only when necessary based on detection confidence levels.
3Reliability
If high-power subsystems are always active to ensure accurate detection, then detection reliability is improved, but power consumption increases
Solution Approach 1:
The system dynamically adjusts processing power based on operational needs rather than maintaining static high-power operation. The cascade architecture enables dynamic activation of higher-power modules only when lower-power modules indicate uncertain or missed detections, optimizing the balance between reliability and power consumption in real-time.
Data Source
AI summary
Systems and methods for identifying audio events in one or more audio streams include the use of a cascade audio spotting system (such as a cascade keyword spotting system (KWS)) to reduce power consumption while maintaining a desired performance. An example cascade audio spotting system may include a first module and a high-power subsystem. The first module is to receive an audio stream from one or more audio streams, process the audio stream to detect a first target sound activity in the audio stream, and provide a first signal in response to detecting the first target sound activity in the audio stream. The high-power subsystem is to (in response to the first signal being provided by the first module) receive the one or more audio streams and process the one or more audio streams to detect a second target sound activity in the one or more audio streams.


