Multi-Level Voice Detection Circuit for Power-Efficient Key Phrase Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing mobile computing devices face challenges in extending battery life and providing seamless voice data capture and processing while maintaining 'always-on' and 'continuously-connected' capabilities, as they consume excessive power during voice recognition tasks.
Innovation Solution
An audio device with a multi-level voice detection circuit system, comprising an acoustic conversion circuit, analog-to-digital converter, and a controller with first, second, and third-level voice detection circuits, which converts acoustic waves into digital data and detects key phrases to efficiently wake up the host device from sleep mode, reducing power consumption and enhancing voice recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If the device maintains 'always-on' capability for prompt voice reaction, then voice recognition responsiveness is improved, but power consumption increases
Solution Approach 1:
The voice detection system is segmented into three hierarchical levels: first-level voice activity detection in analog domain, second-level beginning syllable detection in digital domain, and third-level complete key phrase detection. This segmentation allows the system to quickly filter out non-voice signals at lower levels before consuming more power for detailed analysis, thus maintaining fast response while reducing overall power consumption.
Solution Approach 2:
The system performs partial voice detection actions at different levels rather than complete analysis continuously. The first-level circuit performs partial detection on analog signals, the second-level detects only beginning syllables, and the third-level detects complete key phrases only when triggered. This partial action approach maintains responsiveness while avoiding excessive power consumption from full continuous analysis.
2Adaptability or versatility
If cloud-based information processing is continuously connected, then information accessibility is improved, but battery life decreases
Solution Approach 1:
The audio device performs preliminary voice activity detection and key phrase recognition locally before transmitting data to the cloud. The first-level and second-level detection circuits pre-process analog and digital signals to identify relevant voice inputs, filtering out unnecessary data before cloud transmission. This preliminary action ensures cloud information accessibility is maintained while reducing power consumption from continuous cloud connectivity.
3Reliability
If seamless voice data capture is implemented, then voice interface service quality is improved, but power consumption increases
Solution Approach 1:
Different detection circuits operate with different quality levels appropriate to their function. The first-level analog voice activity detection uses simpler processing suitable for initial filtering, while the second-level digital beginning syllable detection uses more sophisticated analysis only when needed. This local quality differentiation ensures high-quality voice interface service where necessary while reducing power consumption in preliminary stages.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The solution enables reduced power consumption, improved voice command recognition rates, and efficient voice processing, allowing devices to remain 'always-on' while minimizing power usage and ensuring seamless voice interface services.
Implementation Method 1
an acoustic conversion circuit, configured to convert an acoustic wave into an analog audio signal
Data Source
AI summary
An audio device and a method thereof are provided. The method is adopted by an audio device to detect a voice, wherein the audio device is coupled to a host device. The method includes an acoustic conversion circuit converting an acoustic wave into an analog audio signal; an analog-to-digital converter (ADC) converting the analog audio signal into digital audio data; a first-level voice detection circuit detecting voice activity in the analog audio signal; a second-level voice detection circuit detecting a beginning syllable of a key phrase in the digital audio data when the voice activity is detected in the digital audio data; and a third-level voice detection circuit detecting the key phrase from the digital audio data only when the beginning syllable of the key phrase is detected in the digital audio data.


