Multi-modal Interaction Framework for Automated Assistants

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computing devices face challenges in efficiently interacting with automated assistants due to the variety of input and output interfaces, limiting seamless interaction between users and third-party computing services, especially for users with disabilities.

Innovation Solution

A multi-modal interaction framework that allows users to engage with automated assistants using verbal and non-verbal inputs, such as vocal commands and graphical user interface interactions, enabling touchless interaction and efficient traversal of dialog state machines, with a client-server architecture that channels data and commands through a secure server portion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple input and output interfaces are supported across different devices, then accessibility and versatility are improved, but system complexity and integration difficulty increase

Engineering Contradiction:
Improveinterface compatibilityVSAvoidsystem integration
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces an automated assistant as an intermediary layer between users and third-party computing services. This assistant receives various input modalities (verbal, visual, tactile) and translates them into standardized commands that can be processed by different services, thereby simplifying the interaction model without sacrificing versatility.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The automated assistant is designed to handle multiple types of interactions (voice commands, graphical interface manipulation, tactile input) through a single unified system. This multi-functional approach allows the same assistant to work across different devices and services, reducing the need for device-specific implementations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If verbal free form natural language input is used, then ease of operation is improved, but processing time and computational resources increase

Engineering Contradiction:
Improveinput convenienceVSAvoidprocessing time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system employs a hybrid approach where fully verbal commands are processed through natural language understanding, but partially structured inputs (such as selecting from predefined options presented visually) are accepted. This reduces the processing burden while maintaining ease of operation for common tasks.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system pre-processes and structures dialog state machines for common third-party services, so that once a user initiates a conversation, the framework is already prepared to handle the interaction efficiently. This preliminary setup reduces real-time processing requirements.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If touchless interaction is enabled, then accessibility for disabled users is improved, but precision of selection and control decreases

Engineering Contradiction:
ImproveaccessibilityVSAvoidselection accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system provides visual feedback through graphical user interfaces that show users which elements are currently selected or highlighted through voice commands. This feedback loop allows users to verify their selections before confirmation, compensating for the lack of tactile feedback and improving selection accuracy.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system adds a visual dimension to voice-based interactions by displaying graphical representations of the dialog state and available options. This additional dimension allows users to see the consequences of their verbal commands and make more precise selections.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11347801B2Multi-modal interaction between users, automated assistants, and other computing services
Publication Date: 2022.05.31 GOOGLE LLC
  • US11347801B2 patent drawing
  • US11347801B2 patent drawing
  • US11347801B2 patent drawing

AI summary

Techniques are described herein for multi-modal interaction between users, automated assistants, and other computing services. In various implementations, a user may engage with the automated assistant in order to further engage with a third party computing service. In some implementations, the user may advance through dialog state machines associated with third party computing service using both verbal input modalities and input modalities other than verbal modalities, such as visual/tactile modalities.