Voice Activated Device Low-Power Sound Detector Duty Cycle
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-based digital assistants require substantial audio processing and battery power for continuous listening, which leads to high power consumption and reduced device efficiency.
Innovation Solution
A voice-activated device with a human-machine interface consisting essentially of an audio interface is used to interact with a voice-based digital assistant, employing low-power sound detectors and a duty cycle operation to reduce power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If the speech-based service continuously listens for a trigger sound, then the voice-based digital assistant can be activated quickly, but the battery power is consumed rapidly
Solution Approach 1:
The system uses periodic action by implementing a duty cycle where the low-power sound detector alternates between active listening periods and sleep periods. During active periods, the detector monitors for trigger sounds; during sleep periods, it conserves battery power. This periodic operation allows the system to maintain quick activation capability while significantly reducing overall power consumption compared to continuous listening.
2Loss of time
If the main processor remains active for speech processing, then the digital assistant responds quickly to commands, but the power consumption increases substantially
Solution Approach 1:
The system applies segmentation by dividing the processing tasks between a low-power sound detector and the main application processor. The low-power detector handles the initial trigger sound detection and wake word recognition, while the main processor is activated only when needed for full speech processing. This segmentation allows quick response to wake words while minimizing the main processor's active time, thereby reducing overall power consumption.
Solution Approach 2:
The low-power sound detector serves as an intermediary between the microphone and the main application processor. It pre-processes audio signals to detect trigger sounds and wake words, filtering out unnecessary processing requests. This intermediary layer enables the main processor to remain in low-power mode longer while still achieving quick response times when actual speech commands are detected.
3Adaptability or versatility
If the human-machine interface includes multiple components (display, touch, audio), then the device functionality is enhanced, but the power consumption and device size increase
Solution Approach 1:
The system extracts and removes unnecessary interface components (display, touch screen, camera) from the voice-activated device, retaining only the essential audio interface components (microphone, speaker, audio processor). This extraction maintains core voice-based functionality while dramatically reducing power consumption and device size, as the audio interface requires minimal power compared to visual and tactile interfaces.
4Measurement precision
If the sound detector operates continuously at high sensitivity, then trigger sounds are detected accurately, but the power consumption increases
Solution Approach 1:
The system implements dynamics by adjusting the sound detector's operational state based on current needs. The detector dynamically transitions between high-sensitivity active mode (when listening for triggers), moderate-sensitivity standby mode (between triggers), and low-power sleep mode (during extended periods without activity). This dynamic operation maintains high detection accuracy when needed while reducing power consumption during idle periods.
Data Source
AI summary
A voice activated device for interaction with a digital assistant is provided. The device comprises a housing, one or more processors, and memory, the memory coupled to the one or more processors and comprising instructions for automatically identifying and connecting to a digital assistant server. The device further comprises a power supply, a wireless network module, and a human-machine interface. The human-machine interface consists essentially of: at least one speaker, at least one microphone, an ADC coupled to the microphone, a DAC coupled to the at least one speaker, and zero or more additional components selected from the set consisting of: a touch-sensitive surface, one or more cameras, and one or more LEDs. The device is configured to act as an interface for speech communications between the user and a digital assistant of the user on the digital assistant server.


