Dual-Processor Voice Command System to Reduce Power and Latency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice command input devices for computerized devices, such as smartphones and IoT devices, face challenges due to high computational requirements for reliable voice recognition, leading to increased power consumption, waste heat generation, and latency, making them costly and impractical for widespread adoption.
Innovation Solution
A system utilizing two processors: a low-power processor executing a loose voice recognition model to detect a wake word, followed by a more capable processor switching to a high-power mode to verify the wake word using a tighter model, reducing overall power consumption and waste heat while maintaining accurate recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a high-power processor is used for reliable voice recognition, then recognition accuracy is improved, but power consumption and waste heat increase
Solution Approach 1:
The voice recognition system is divided into two separate processors: a low-power processor that handles basic wake word detection, and a high-power processor that performs accurate voice command recognition. This segmentation allows each processor to be optimized for its specific function, reducing overall power consumption while maintaining recognition accuracy.
Solution Approach 2:
The system dynamically switches between different processing modes based on the operational phase. The low-power processor operates continuously for wake word detection, while the high-power processor is activated only when a wake word is detected. This dynamic operation reduces average power consumption compared to continuously running the high-power processor.
2Reliability
If a high-power processor is used for reliable voice recognition, then recognition accuracy is improved, but waste heat generation increases
Solution Approach 1:
By segmenting the voice recognition tasks between two processors, the system reduces waste heat generation. The low-power processor generates minimal heat during continuous operation, and the high-power processor generates heat only briefly when activated for accurate recognition, rather than continuously generating heat.
Solution Approach 2:
The high-power processor operates periodically rather than continuously - it is activated only when the low-power processor detects a wake word. This periodic operation significantly reduces cumulative waste heat generation while maintaining the necessary recognition accuracy.
3Reliability
If a high-power processor is used for reliable voice recognition, then recognition reliability is improved, but device cost increases
Solution Approach 1:
The system uses a low-power processor for continuous wake word detection and a high-power processor only when needed for accurate recognition. This segmentation allows the use of a cheaper low-power processor for the majority of operation time, reducing overall device cost while maintaining recognition reliability through the high-power processor's periodic intervention.
Solution Approach 2:
The low-power processor acts as an intermediary that filters incoming audio signals, activating the expensive high-power processor only when necessary. This intermediary approach reduces the burden on the high-power processor and allows the system to achieve reliable recognition with lower average computational cost.
4Measurement precision
If continuous high-power processing is used for wake word detection, then detection accuracy is improved, but latency increases
Solution Approach 1:
The detection process is segmented into two stages: rapid wake word detection by the low-power processor, followed by accurate verification by the high-power processor only when needed. This segmentation reduces latency by handling the common case (wake word detection) with the faster low-power processor.
Solution Approach 2:
The low-power processor performs preliminary wake word detection continuously at low power, preparing the system for potential high-power processing. This preliminary action reduces latency by having the detection mechanism already active and ready, rather than waiting to activate high-power processing.
Data Source
AI summary
A system including at least one computerized device with voice command capability processed remotely includes a low power processor, executing a loose algorithmic model to recognize a wake word prefix in a voice command, the loose model having a low false rejection rate but suffering a high false acceptance rate, and a second processor which can operate in at least a low power/low clock rate mode and a high power/high clock rate mode. When the first processor determines the presence of the wake word, it causes the second processor to switch to the high power/high clock rate mode and to execute a tight algorithmic model to verify the presence of the wake word. By using the two processors in this manner, the average overall power required by the computerized device is reduced, as is the amount of waste heat generated by the system.


