Dynamic Speech Recognition Power Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice assistant technologies in electronic devices have poor sensitivity and high power consumption due to frequent wake-ups for voice detection, which affects user experience and energy efficiency.
Innovation Solution
A dynamic speech recognition method that uses a digital microphone, processing circuits, and memory access circuits to selectively perform stages of sound data processing based on power budget and recognition interval time, optimizing power consumption and sensitivity by determining whether to perform a DMA stage or a speech recognition stage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the voice assistant wakes up regularly to detect human voice, then the sensitivity of voice detection is improved, but the power consumption increases
Solution Approach 1:
The speech recognition process is divided into multiple stages: a first stage with lower power consumption for basic voice activity detection, and a second stage with higher power consumption for full speech recognition. The system dynamically transitions between stages based on voice detection results and power budget conditions, allowing sensitive detection only when power is available.
Solution Approach 2:
The system dynamically adjusts its operating mode between different recognition stages based on real-time conditions. When power budget allows and voice activity is detected, the system transitions to the second stage for more sensitive and accurate speech recognition. Otherwise, it remains in the first stage to conserve power, creating a dynamic balance between sensitivity and energy consumption.
2Measurement precision
If the voice assistant performs full speech recognition frequently, then the recognition accuracy is improved, but the average power consumption increases
Solution Approach 1:
Instead of performing full speech recognition continuously, the system performs partial action by first conducting voice activity detection in the first stage. Only when voice activity is detected and power budget permits does the system proceed to the more resource-intensive second stage for accurate speech recognition, avoiding excessive power consumption from unnecessary full recognition operations.
Solution Approach 2:
The system changes operational parameters dynamically by adjusting which processing stage is active based on power budget conditions and voice detection results. This parameter change allows the system to optimize between recognition accuracy and power consumption by selecting the appropriate level of processing intensity for each situation.
3Use of energy by moving object
If the system reduces wake-up frequency to save power, then the power consumption is reduced, but the voice detection sensitivity deteriorates
Solution Approach 1:
The system performs preliminary voice activity detection in the first stage before committing to full speech recognition in the second stage. This preliminary action allows the system to maintain a form of continuous monitoring with low power consumption, ready to transition to high-sensitivity mode when voice activity is detected and power is available, thus maintaining sensitivity without continuous high-power operation.
Data Source
AI summary
A dynamic speech recognition method includes performing a first stage: detecting sound data by using a digital microphone and storing the sound data in a first memory, generating a human voice detection signal in response to detecting a human voice from the sound data, and determining to selectively perform a second stage or a third stage according to a total effective data volume, a transmission bit rate of the digital microphone and a recognition interval time. In the second stage, the first processing circuit outputs a first command to a second processing circuit, and the second processing circuit instructs a memory access circuit to operate. In the third stage, the first processing circuit outputs a second command to the second processing circuit, and the second processing circuit instructs the memory access circuit to operate, and the second processing circuit determines whether the speech data matches a predetermined speech command.


