Voice Keyword Recognition Using Multi-Stage DSP and Server Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Electronic apparatuses often fail to accurately process user utterances due to incorrect recognition of keywords, leading to unintended operation or failure to operate as intended.
Innovation Solution
An electronic apparatus with a digital signal processing circuit, processor, and memory that receives user utterance data, determines the presence of a specified word, and communicates with an external server to verify the recognition, ensuring accurate activation of voice-based input systems based on the determination.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If the electronic apparatus uses a simple keyword recognition method, then the processing speed is fast, but the recognition accuracy deteriorates leading to incorrect operation
Solution Approach 1:
The keyword recognition process is segmented into multiple independent determination stages: first determination by DSP circuit, second determination by processor, and third determination by external server. Each stage processes the user utterance independently and contributes to the final recognition result, allowing the system to maintain fast local processing while incorporating accurate remote verification.
Solution Approach 2:
The system implements feedback mechanisms where the external server provides verification results back to the processor, which then adjusts the activation decision. The processor receives text generated from the first data by the external server and uses this feedback to make the final determination on whether to activate the voice-based input system, improving overall recognition accuracy.
2Measurement precision
If the electronic apparatus performs multiple determination steps, then the recognition accuracy is improved, but the device complexity increases
Solution Approach 1:
The complex recognition system is segmented into distinct functional modules: DSP circuit for first determination, processor for second determination and coordination, and external server for third determination. This segmentation allows each component to have a specific, simplified function while the overall system achieves high accuracy through their coordinated operation.
Solution Approach 2:
The processor acts as an intermediary that coordinates between the DSP circuit and the external server. It receives first data from the DSP circuit, manages the communication with the external server, and integrates the results from multiple determination steps, thereby managing system complexity while maintaining high recognition accuracy.
3Ease of operation
If the electronic apparatus activates the voice-based input system frequently, then the user responsiveness is high, but the energy consumption increases
Solution Approach 1:
The system performs preliminary determination steps using the DSP circuit and processor before fully activating the voice-based input system. By conducting first and second determinations locally before engaging the external server and final activation, the system prepares in advance and only activates the full system when necessary, improving responsiveness while avoiding unnecessary energy consumption.
Solution Approach 2:
The system performs partial determination actions (first and second determinations) that are sufficient for preliminary assessment but not complete activation. This partial action allows the system to evaluate user input quickly and only commit to full activation when the preliminary checks indicate a genuine need, balancing responsiveness with energy efficiency.
Data Source
AI summary
An apparatus comprising one or more processors, a communication circuit, and a memory for storing instructions, which when executed, performs a method of recognizing a user utterance. The method comprises: receiving first data associated with a user utterance, performing, a first determination to determine whether the user utterance includes the first data and a specified word, performing a second determination to determine whether the first data includes the specified word, transmitting the first data to an external server, receiving a text generated from the first data by the external server, performing a third determination to determine whether the received text matches the specified word, and determining whether to activate the voice-based input system based on the third determination.


