Low-Power Voice Trigger With Segmented Sound Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-based digital assistants require tactile input to activate, consuming power resources and detracting from a hands-free experience, and continuous audio processing for voice triggers is power-intensive.
Innovation Solution
Implement a low-power voice trigger system using sound detectors that recognize specific words or phrases without requiring continuous audio processing, employing a duty cycle and multiple detectors to minimize power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If continuous audio processing is used for voice trigger, then voice activation reliability is improved, but power consumption increases
Solution Approach 1:
The voice trigger system is segmented into multiple detectors operating at different processing levels. A first sound detector performs basic audio analysis, and only when triggered does a second sound detector perform more comprehensive analysis. This segmentation allows the system to maintain reliability through multi-level verification while dramatically reducing power consumption by keeping most processing inactive.
Solution Approach 2:
The sound detectors operate periodically rather than continuously, using a duty cycle approach where they are activated only at specific intervals or when preliminary conditions are met. The first sound detector monitors audio continuously at low power, but the second sound detector and full processing only activate periodically when the first detector identifies potential triggers, balancing reliability with power savings.
2Measurement precision
If multiple sound detectors are used to improve detection accuracy, then voice recognition precision is improved, but device complexity increases
Solution Approach 1:
The detection system is segmented into a first sound detector and a second sound detector, each performing specific functions. The first detector handles preliminary audio analysis with lower computational requirements, while the second detector performs more sophisticated analysis only when needed. This segmentation improves precision through specialized detection layers while managing complexity by dividing rather than consolidating functions.
Solution Approach 2:
The system applies partial action by using the first sound detector for continuous monitoring with limited processing, and only deploying the full second sound detector analysis when the first detector identifies potential voice triggers. This partial deployment of detection capabilities maintains high precision for actual voice commands while avoiding the complexity of always-running full analysis.
3Ease of operation
If audio monitoring is activated for hands-free operation, then ease of operation is improved, but energy consumption increases
Solution Approach 1:
The audio monitoring system uses periodic action by activating the first sound detector continuously at low power for hands-free operation, but only activating the second sound detector and full processing periodically when preliminary audio analysis identifies potential triggers. This maintains ease of hands-free operation while minimizing battery power consumption through duty-cycled processing.
Solution Approach 2:
The first sound detector performs preliminary audio analysis continuously to identify potential voice triggers, preparing the system for hands-free operation. Only when this preliminary action detects something of interest does the system activate more comprehensive processing. This preliminary action enables ease of operation while controlling energy consumption by avoiding full processing until necessary.
Data Source
AI summary
A method for operating a voice trigger is provided. In some implementations, the method is performed at an electronic device including one or more processors and memory storing instructions for execution by the one or more processors. The method includes receiving a sound input. The sound input may correspond to a spoken word or phrase, or a portion thereof. The method includes determining whether at least a portion of the sound input corresponds to a predetermined type of sound, such as a human voice. The method includes, upon a determination that at least a portion of the sound input corresponds to the predetermined type, determining whether the sound input includes predetermined content, such as a predetermined trigger word or phrase. The method also includes, upon a determination that the sound input includes the predetermined content, initiating a speech-based service, such as a voice-based digital assistant.


