Cascade Audio Spotting System Sensitivity Mode
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Audio controlled devices face challenges in reducing power consumption while maintaining effective detection of spoken keywords or audio events, leading to increased false acceptances and rejections due to the always-on nature of speech recognition systems, which results in inefficient power usage and user experience issues.
Innovation Solution
A cascade audio spotting system is implemented, comprising multiple modules that operate sequentially, with lower power-consuming modules detecting initial audio activities triggering higher power modules only when necessary, reducing overall power consumption without compromising performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the audio processing portion operates in always-on mode to detect keywords or audio events, then the detection reliability is improved, but the power consumption increases
Solution Approach 1:
The audio processing system is divided into multiple modules with different power consumption levels and detection capabilities. A low-power module continuously monitors for basic audio events, while high-power modules are activated only when needed, segmenting the always-on function into hierarchical layers that balance detection reliability with power efficiency
Solution Approach 2:
The low-power module performs preliminary detection of audio events before activating the high-power module. This preliminary action filters out non-critical events, ensuring that the high-power module only processes potentially relevant audio streams, thereby maintaining detection reliability while reducing overall power consumption
2Reliability
If the audio processing portion is activated continuously, then false acceptances are reduced, but power consumption increases
Solution Approach 1:
The system dynamically adjusts its processing power based on detected audio conditions. The low-power module continuously adapts its detection sensitivity and triggers the high-power module only when confidence thresholds are met or ambiguous events occur, dynamically balancing false acceptance rates with power consumption
3Reliability
If the high-power subsystem is activated frequently, then detection performance is improved, but battery life is reduced
Solution Approach 1:
The low-power module serves as an intermediary between the continuous audio input and the high-power subsystem. It pre-processes and filters audio streams, acting as a gatekeeper that activates the high-power module only when necessary, thereby extending battery life while maintaining detection performance through selective activation
Data Source
AI summary
An audio spotting system configured for various operating modes including a regular mode and sensitivity mode is described. An example cascade audio spotting system may include a high-power subsystem including a high-power trigger and a transfer module. This high-power trigger includes one or more detection models used to detect whether a target sound activity is included in the one or more audio streams. The one or more detection models are associated with a first set of hyperparameters when the cascade audio spotting system is in a regular mode, and the one or more detection models are associated with a second set of hyperparameters when the cascade audio spotting system is in a sensitivity mode. The transfer module provides at least one of one or more processed audio streams for further processing in response to the high-power trigger detecting the target sound activity in the one or more audio streams.


