Multi-Modal Voice Commands with Affordance-Guided Task Execution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional digital assistants face challenges in performing complex tasks due to the limitations of speech input, such as requiring precise parameters and handling ambiguous natural language, leading to a poor user experience, especially for inexperienced users.

Innovation Solution

Implementing multi-modal inputs that combine activation of an affordance on a device with user speech to determine user intent, allowing for more precise task execution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speech input alone is used for digital assistant commands, then the interface is simple and easy to use, but the ability to perform complex tasks with precise parameters is insufficient

Engineering Contradiction:
Improveease of useVSAvoidtask execution capability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent combines speech input with graphical user interface elements (buttons, sliders, keyboards) into a unified interaction system. The digital assistant can process commands through multiple modalities simultaneously or sequentially, allowing users to switch between speech and graphical controls based on task complexity, thereby resolving the contradiction between ease of use and task execution capability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system dynamically adapts its input requirements based on the detected task type and user intent. For simple tasks, speech alone suffices; for complex tasks requiring precise parameters, the system dynamically introduces graphical controls. This dynamic adjustment of interaction complexity allows the system to maintain ease of use for simple operations while gaining versatility for complex tasks.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If speech input is used for complex tasks requiring precise parameters, then the interface remains simple, but the precision and accuracy of parameter specification deteriorates

Engineering Contradiction:
Improveparameter specification precisionVSAvoidinput interface complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary analysis of the user's speech command to detect task type and required parameters before full processing. This preliminary action allows the system to prepare appropriate graphical controls in advance, presenting only the necessary controls for the specific task, thereby maintaining precision without unnecessarily increasing overall interface complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of providing all possible controls simultaneously, the system applies local quality by presenting only the specific graphical controls relevant to the detected task and parameters. For example, if a music playback task is detected, only playback controls are shown, not all device controls. This selective presentation maintains precision for relevant parameters while minimizing perceived complexity.

Inventive Principle:
Principle #3Local quality

3Reliability

If multi-modal inputs combining affordance activation and speech are implemented, then task execution accuracy improves, but the system complexity increases

Engineering Contradiction:
Improvetask completion accuracyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary processing layer that receives both speech input and graphical control inputs, integrates them, and determines user intent. This intermediary layer acts as a mediator between the multiple input modalities and the task execution engine, coordinating the different inputs and resolving conflicts or ambiguities, thereby improving reliability while managing system complexity through structured integration.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Ease of manufacture

If conventional speech-only processing is used, then the system is easy to implement, but the handling of ambiguous natural language terms fails

Engineering Contradiction:
Improvesystem implementation easeVSAvoidambiguous term resolution
Core Design Contradiction:
Ease of manufactureVSLoss of information

Solution Approach 1:

The system implements feedback mechanisms where the digital assistant can request clarification or additional input when ambiguous terms are detected. The system processes the initial speech input, identifies ambiguities, and feeds back to the user with specific questions or alternative interpretations, allowing the user to refine their input. This feedback loop resolves ambiguous terms while maintaining relatively simple implementation through structured dialogue management.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12367879B2Multi-modal inputs for voice commands
Publication Date: 2025.07.22 APPLE INC
  • US12367879B2 patent drawing
  • US12367879B2 patent drawing
  • US12367879B2 patent drawing

AI summary

Systems and processes for operating an intelligent automated assistant are provided. In one example process, a first input including activation of an affordance is received. A domain associated with the affordance is determined. A second input including user speech is received, where a user intent is determined based on the domain and the user speech. A determination is made whether the user intent includes a command associated with the affordance. In accordance with a determination that the user intent includes a command associated with the affordance, a task in furtherance of the command is performed.