Multi-Stage Target Sound Detection for Lower Power Audio Monitoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Always-on audio context detection systems in electronic devices result in high power consumption, reducing battery life, especially in mobile devices, and increasing system complexity due to the need to detect multiple sound events.
Innovation Solution
A multi-stage target sound detector with a binary classification first stage and a more powerful second stage, where the first stage reduces power consumption by performing binary classification and activating the second stage only when target sounds are detected, allowing for high-performance sound classification with reduced average power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If an always-on audio context detection system is used to detect multiple sound events, then sound detection capability is improved, but power consumption increases
Solution Approach 1:
The audio context detection system is divided into two distinct stages: a first stage that performs lightweight binary classification to detect target sounds, and a second stage that performs more complex multiple sound event classification. This segmentation allows the system to maintain high sound detection capability while reducing power consumption by avoiding continuous execution of complex processing.
Solution Approach 2:
The system dynamically adjusts its processing intensity based on detected conditions. The first stage continuously monitors audio for target sounds using low-power binary classification, and only activates the second stage with multiple sound event detection when target sounds are detected. This dynamic adaptation enables the system to maintain versatility when needed while conserving power during normal operation.
2Adaptability or versatility
If the number of sound events to be detected is increased, then sound detection capability is improved, but system complexity increases
Solution Approach 1:
The classification task is segmented into two stages with different complexity levels. The first stage handles binary classification (target sound vs. non-target sound) with simpler processing, while the second stage handles multiple sound event classification with more complex processing. This segmentation reduces overall system complexity by avoiding continuous execution of complex algorithms.
Solution Approach 2:
The system applies partial action by using simpler binary classification in the first stage to screen audio data, and only applies the more complex multiple sound event classification in the second stage when necessary (when target sounds are detected). This partial application of complex processing reduces system operational complexity while maintaining detection capability.
3Reliability
If continuous audio processing is performed to detect sound events, then detection reliability is improved, but power consumption increases
Solution Approach 1:
The first stage performs preliminary binary classification of audio data to identify potential target sounds before activating the second stage for detailed analysis. This preliminary action ensures detection reliability by maintaining continuous monitoring capability while reducing power consumption by avoiding continuous execution of complex processing algorithms.
Solution Approach 2:
The system employs periodic action through the sequential activation of processing stages. The first stage operates continuously with low power consumption, and the second stage is periodically activated only when the first stage detects target sounds. This periodic activation pattern maintains detection reliability while significantly reducing average power consumption.
Data Source
AI summary
A device to perform target sound detection includes a memory including a buffer configured to store audio data. The device includes one or more processors coupled to the memory. The one or more processors are configured to receive the audio data from the buffer. The one or more processors are configured to detect the presence or absence of one or more target non-speech sounds in the audio data. The one or more processors are further configured to generate a user interface signal, to indicate one of the one or more target non-speech sounds has been detected, and provide the user interface signal to an output device. The device further comprises the output device that is configured to output a visual representation associated with the one of the one or more target non-speech sounds has been detected.


