Voice Assistant Intent Clarification via External Server Mediation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current electronic devices face challenges in accurately interpreting and executing tasks based on ambiguous user utterances, as they often struggle to understand the intent behind vague voice inputs, leading to inefficiencies in performing requested actions.

Innovation Solution

An electronic device equipped with a processor, memory, touchscreen display, microphone, and wireless communication circuit that transmits user utterances to an external server for processing, receives path rules, and displays sample utterances for user selection, allowing the device to perform tasks aligned with the user's intent.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the device processes ambiguous user utterances directly without external validation, then the response speed is faster, but the accuracy of task execution deteriorates

Engineering Contradiction:
Improveaccuracy of intent recognitionVSAvoidresponse time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces an external server as an intermediary component that receives ambiguous user utterances from the electronic device, processes them using advanced natural language understanding algorithms, and returns clarified intent information. This mediator handles the complex processing task that would otherwise slow down the device's response time while improving accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary processing of user utterances by transmitting them to an external server before final task execution. The server pre-processes and validates the utterance intent, returning structured information that the device can quickly act upon, thus improving both accuracy and maintaining response speed.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If the device displays multiple sample utterances for user selection, then the accuracy of task execution improves, but the ease of operation deteriorates

Engineering Contradiction:
Improveaccuracy of intent recognitionVSAvoiduser interaction complexity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

Instead of presenting all possible sample utterances to the user, the system displays only a subset of the most relevant and probable options based on the analyzed input. This partial presentation approach maintains accuracy by showing targeted options while reducing user interaction complexity by limiting the choice set to manageable proportions.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If the device processes all user utterances locally without external server assistance, then the device complexity is reduced, but the reliability of task execution deteriorates

Engineering Contradiction:
Improvereliability of task executionVSAvoidsystem architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent employs an external server as a mediator that handles complex natural language understanding and intent recognition tasks. This separates the reliability-critical processing functions from the device itself, allowing the device to maintain simpler architecture while achieving higher reliability through the server's advanced processing capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11170768B2Device for performing task corresponding to user utterance
Publication Date: 2021.11.09 SAMSUNG ELECTRONICS CO LTD
  • US11170768B2 patent drawing
  • US11170768B2 patent drawing
  • US11170768B2 patent drawing

AI summary

An electronic device includes a touchscreen display, a microphone, at least one speaker, a processor and a memory which stores instructions that cause the processor to receive a user utterance including a request for performing a task with the electronic device, to transmit data associated with the user utterance to an external server, to receive a response from the external server including sample utterances representative of an intent of the user utterance and the sample utterances being selected by the external server based on the user utterance, to display the sample utterances on the touchscreen display, to receive a user input to select one of the sample utterances, and to perform the task by causing the electronic device to follow a sequence of states associated with the selected one of the sample utterances.