Two-Stage Voice Command Processing to Reduce Power and Heat

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice command systems for computerized devices, such as smartphones and IoT devices, face challenges due to high computational requirements for reliable voice recognition, leading to increased power consumption, heat generation, and latency, making them costly and impractical for devices with limited resources like batteries or HVAC controllers.

Innovation Solution

A two-stage voice recognition system using a low-power processor for initial wake word detection with a loose algorithm and a secondary processor for verification with a tighter algorithm, reducing overall power consumption and heat generation while maintaining low false rejection and acceptance rates.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a high-powered processor is used for reliable voice recognition, then voice recognition accuracy is improved, but power consumption and heat generation increase

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidpower consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The voice recognition system is segmented into two distinct processing stages: a first processor handles initial wake word detection with lower computational requirements, while a second processor performs more intensive verification only when the wake word is detected. This segmentation allows the system to maintain high recognition accuracy while reducing average power consumption, as the high-powered second processor operates only intermittently rather than continuously.

Inventive Principle:
Principle #1Segmentation

2Reliability

If a high-powered processor is used for reliable voice recognition, then voice recognition accuracy is improved, but heat generation increases

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidheat generation
Core Design Contradiction:
ReliabilityVSTemperature

Solution Approach 1:

The processing workload is segmented between two processors with different computational capabilities. The first processor handles the majority of processing tasks at lower power levels, generating minimal heat. The second processor, which generates more heat, is activated only when necessary for verification, thereby reducing overall heat generation while maintaining recognition accuracy.

Inventive Principle:
Principle #1Segmentation

3Reliability

If a high-powered processor is used for reliable voice recognition, then voice recognition accuracy is improved, but device cost increases

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoiddevice cost
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The system uses two processors with different capability levels rather than a single high-powered processor. The first processor can be a lower-cost, lower-power component suitable for basic detection tasks. The second processor, which is more capable and expensive, is used only intermittently for verification. This segmentation allows the system to achieve high recognition accuracy while reducing the overall cost compared to using a high-powered processor continuously.

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If continuous processing is used for wake word detection, then detection accuracy is improved, but latency increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidprocessing latency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The first processor performs preliminary wake word detection continuously at lower computational levels. When a potential wake word is detected, the system prepares the second processor for verification. This preliminary action allows the system to maintain high detection accuracy while reducing latency, as the verification process is initiated only when necessary rather than processing every audio sample at full computational power.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11521614B2Device with voice command input capabtility
Publication Date: 2022.12.06 GENERAC POWER SYSTEMS INC
  • US11521614B2 patent drawing
  • US11521614B2 patent drawing
  • US11521614B2 patent drawing

AI summary

A system including at least one computerized device with voice command capability processed remotely includes a low power processor, executing a loose algorithmic model to recognize a wake word prefix in a voice command, the loose model having a low false rejection rate but suffering a high false acceptance rate, and a second processor which can operate in at least a low power/low clock rate mode and a high power/high clock rate mode. When the first processor determines the presence of the wake word, it causes the second processor to switch to the high power/high clock rate mode and to execute a tight algorithmic model to verify the presence of the wake word. By using the two processors in this manner, the average overall power required by the computerized device is reduced, as is the amount of waste heat generated by the system.