Intelligent Assistance Utterance Processing via Intent Masking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current intelligent assistance services face limitations in utterance recognition performance and accuracy, leading to increased communication usage and slower response times due to the need for data transmission to servers for precise recognition, which degrades user satisfaction.
Innovation Solution
An electronic device and server system that uses intent masking information to determine whether to process user utterances locally or remotely, leveraging both device and server speech processing capabilities to optimize recognition and reduce communication overhead.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the electronic device transmits utterance data to a server for recognition, then utterance recognition accuracy is improved, but communication usage increases and response time becomes slower
Solution Approach 1:
The system segments utterance processing into two parts: simple utterances are processed locally by the electronic device's speech processing module, while complex utterances are transmitted to the server for processing. This segmentation allows the system to respond quickly to simple commands while maintaining high accuracy for complex recognition tasks.
Solution Approach 2:
The server acts as an intermediary that receives utterance data from the electronic device, processes it using its speech processing module, and returns recognition results. This intermediary approach enables the electronic device to leverage server-side processing power for accurate recognition without permanently increasing communication overhead.
2Measurement precision
If the electronic device transmits utterance data to a server for recognition, then utterance recognition accuracy is improved, but communication usage increases
Solution Approach 1:
The system segments utterance processing into local and remote components. The speech processing module at the electronic device handles simple utterances locally, avoiding unnecessary communication. Only complex utterances that exceed local processing capabilities are transmitted to the server, thereby reducing overall communication usage while maintaining recognition accuracy when needed.
Solution Approach 2:
Instead of transmitting all utterances to the server, the system applies partial action by selectively transmitting only those utterances that require server-side processing. This approach achieves sufficient recognition accuracy for complex cases while avoiding the excessive communication overhead of universal server transmission.
3Speed
If the electronic device uses only its own speech processing module, then response time is reduced, but utterance recognition accuracy is limited
Solution Approach 1:
The system segments processing responsibilities between the electronic device and server based on utterance complexity. Simple utterances are processed locally for fast response, while complex utterances are routed to the server for accurate recognition. This segmentation allows the system to optimize for both speed and accuracy depending on the specific utterance.
Solution Approach 2:
The system dynamically determines whether to process an utterance locally or transmit it to the server based on real-time assessment of utterance complexity and confidence levels. This dynamic approach enables the system to adapt its processing strategy, achieving fast response times for simple cases while ensuring high accuracy for complex cases.
Data Source
AI summary
An electronic device includes at least one communication circuit, at least one microphone, at least one processor operatively connected to the at least one communication circuit and the at least one microphone, and at least one memory operatively connected to the at least one processor. The at least one memory is configured to store instructions. The at least one processor is configured to store intent masking information that defines an utterance processing target for at least one intent, in the memory. When an utterance indicating a speech based intelligent assistance service through the at least one microphone is received, the at least one processor is configured to determine that a processing target of the received utterance is one of the electronic device or a server connected through the at least one communication circuit, based on the intent masking information.


