Multi-modal Interaction Framework for Automated Assistants
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computing devices face challenges in efficiently interacting with automated assistants due to the variety of input and output interfaces, limiting seamless interaction between users and third-party computing services, especially for users with disabilities.
Innovation Solution
A multi-modal interaction framework that allows users to engage with automated assistants using verbal and non-verbal inputs, such as vocal commands and graphical user interface interactions, enabling touchless interaction and efficient traversal of dialog state machines, with a client-server architecture that channels data and commands through a secure server portion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple input and output interfaces are supported across different devices, then accessibility and versatility are improved, but system complexity and integration difficulty increase
Solution Approach 1:
The patent introduces an automated assistant as an intermediary layer between users and third-party computing services. This assistant receives various input modalities (verbal, visual, tactile) and translates them into standardized commands that can be processed by different services, thereby simplifying the interaction model without sacrificing versatility.
Solution Approach 2:
The automated assistant is designed to handle multiple types of interactions (voice commands, graphical interface manipulation, tactile input) through a single unified system. This multi-functional approach allows the same assistant to work across different devices and services, reducing the need for device-specific implementations.
2Ease of operation
If verbal free form natural language input is used, then ease of operation is improved, but processing time and computational resources increase
Solution Approach 1:
The system employs a hybrid approach where fully verbal commands are processed through natural language understanding, but partially structured inputs (such as selecting from predefined options presented visually) are accepted. This reduces the processing burden while maintaining ease of operation for common tasks.
Solution Approach 2:
The system pre-processes and structures dialog state machines for common third-party services, so that once a user initiates a conversation, the framework is already prepared to handle the interaction efficiently. This preliminary setup reduces real-time processing requirements.
3Adaptability or versatility
If touchless interaction is enabled, then accessibility for disabled users is improved, but precision of selection and control decreases
Solution Approach 1:
The system provides visual feedback through graphical user interfaces that show users which elements are currently selected or highlighted through voice commands. This feedback loop allows users to verify their selections before confirmation, compensating for the lack of tactile feedback and improving selection accuracy.
Solution Approach 2:
The system adds a visual dimension to voice-based interactions by displaying graphical representations of the dialog state and available options. This additional dimension allows users to see the consequences of their verbal commands and make more precise selections.
Data Source
AI summary
Techniques are described herein for multi-modal interaction between users, automated assistants, and other computing services. In various implementations, a user may engage with the automated assistant in order to further engage with a third party computing service. In some implementations, the user may advance through dialog state machines associated with third party computing service using both verbal input modalities and input modalities other than verbal modalities, such as visual/tactile modalities.


