Multi-Modal Voice Commands with Affordance-Guided Task Execution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional digital assistants face challenges in performing complex tasks due to the limitations of speech input, such as requiring precise parameters and handling ambiguous natural language, leading to a poor user experience, especially for inexperienced users.
Innovation Solution
Implementing multi-modal inputs that combine activation of an affordance on a device with user speech to determine user intent, allowing for more precise task execution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speech input alone is used for digital assistant commands, then the interface is simple and easy to use, but the ability to perform complex tasks with precise parameters is insufficient
Solution Approach 1:
The patent combines speech input with graphical user interface elements (buttons, sliders, keyboards) into a unified interaction system. The digital assistant can process commands through multiple modalities simultaneously or sequentially, allowing users to switch between speech and graphical controls based on task complexity, thereby resolving the contradiction between ease of use and task execution capability.
Solution Approach 2:
The system dynamically adapts its input requirements based on the detected task type and user intent. For simple tasks, speech alone suffices; for complex tasks requiring precise parameters, the system dynamically introduces graphical controls. This dynamic adjustment of interaction complexity allows the system to maintain ease of use for simple operations while gaining versatility for complex tasks.
2Measurement precision
If speech input is used for complex tasks requiring precise parameters, then the interface remains simple, but the precision and accuracy of parameter specification deteriorates
Solution Approach 1:
The system performs preliminary analysis of the user's speech command to detect task type and required parameters before full processing. This preliminary action allows the system to prepare appropriate graphical controls in advance, presenting only the necessary controls for the specific task, thereby maintaining precision without unnecessarily increasing overall interface complexity.
Solution Approach 2:
Instead of providing all possible controls simultaneously, the system applies local quality by presenting only the specific graphical controls relevant to the detected task and parameters. For example, if a music playback task is detected, only playback controls are shown, not all device controls. This selective presentation maintains precision for relevant parameters while minimizing perceived complexity.
3Reliability
If multi-modal inputs combining affordance activation and speech are implemented, then task execution accuracy improves, but the system complexity increases
Solution Approach 1:
The patent introduces an intermediary processing layer that receives both speech input and graphical control inputs, integrates them, and determines user intent. This intermediary layer acts as a mediator between the multiple input modalities and the task execution engine, coordinating the different inputs and resolving conflicts or ambiguities, thereby improving reliability while managing system complexity through structured integration.
4Ease of manufacture
If conventional speech-only processing is used, then the system is easy to implement, but the handling of ambiguous natural language terms fails
Solution Approach 1:
The system implements feedback mechanisms where the digital assistant can request clarification or additional input when ambiguous terms are detected. The system processes the initial speech input, identifies ambiguities, and feeds back to the user with specific questions or alternative interpretations, allowing the user to refine their input. This feedback loop resolves ambiguous terms while maintaining relatively simple implementation through structured dialogue management.
Data Source
AI summary
Systems and processes for operating an intelligent automated assistant are provided. In one example process, a first input including activation of an affordance is received. A domain associated with the affordance is determined. A second input including user speech is received, where a user intent is determined based on the domain and the user speech. A determination is made whether the user intent includes a command associated with the affordance. In accordance with a determination that the user intent includes a command associated with the affordance, a task in furtherance of the command is performed.


