Multimodal Response Generation for Dynamic Modality Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing multimodal client devices face inefficiencies in generating output due to the need for separate responses tailored to each modality, leading to increased memory storage and computational requirements, as well as latency issues when switching between interaction types.

Innovation Solution

Implementing a dynamic generation of client device output using a single multimodal response that adapts to the current modality of the device, determined by sensor data and user interface input, allowing for the selection of appropriate components to render output efficiently across various modalities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If separate responses are generated for each modality, then output accuracy for each specific modality is improved, but memory storage and computational requirements increase

Engineering Contradiction:
Improveoutput accuracyVSAvoidmemory storage requirements
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent merges multiple modality-specific responses into a single multimodal response that contains all possible output components (audio, visual, haptic, etc.). Instead of storing separate responses for each modality combination, the system stores one comprehensive response that can be dynamically adapted to any current modality by selecting and rendering only the relevant components.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The single multimodal response is designed to serve multiple modalities simultaneously. It contains universal output components that can be rendered across different modalities (audio output, visual output, haptic output), making the response structure multi-functional and adaptable to any current modality without requiring separate specialized responses.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If separate responses are generated for each modality, then modality-specific output quality is improved, but computational requirements increase

Engineering Contradiction:
Improvemodality-specific output qualityVSAvoidcomputational requirements
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent combines multiple modality-specific processing into a single multimodal response generation process. The system processes user input once and generates a comprehensive response containing all modality components, eliminating the need for separate processing pipelines for each modality and reducing overall computational requirements.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The multimodal response structure is designed to be universally applicable across all modalities. A single response generation mechanism produces output that can be rendered in any current modality, making the computational process multi-functional and eliminating redundant calculations that would occur with separate modality-specific responses.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If separate responses are generated for each modality, then response accuracy for each interaction type is improved, but latency increases when switching between interaction types

Engineering Contradiction:
Improveresponse accuracyVSAvoidlatency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by generating a complete multimodal response that includes all possible output components before the system needs to switch between interaction types. The response is prepared in advance with all modality components (audio, visual, haptic) already processed and ready, so when the current modality changes, the system can immediately select and render the appropriate components without waiting for additional processing or requests.

Inventive Principle:
Principle #10Preliminary action

4Quantity of substance

If a single multimodal response is used, then memory storage and computational efficiency are improved, but the complexity of selecting appropriate components increases

Engineering Contradiction:
Improvememory storage efficiencyVSAvoidcomponent selection complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent segments the single multimodal response into distinct, independently selectable components corresponding to different output modalities (audio output component, visual output component, haptic output component). This segmentation allows the system to easily select and render only the relevant components based on the current modality, reducing the complexity of component selection while maintaining storage efficiency.

Inventive Principle:
Principle #1Segmentation

5Power

If a single multimodal response is used, then computational efficiency is improved, but the complexity of adapting to different modalities increases

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidmodality adaptation complexity
Core Design Contradiction:
PowerVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamics by making the multimodal response structure adaptable to different current modalities. The system dynamically selects which output components to render based on the detected current modality (voice only, visual only, multimodal interaction, etc.), allowing the single response to efficiently adapt to various interaction types without requiring separate static responses for each modality.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11164576B2Multimodal responses
Publication Date: 2021.11.02 GOOGLE LLC
  • US11164576B2 patent drawing
  • US11164576B2 patent drawing
  • US11164576B2 patent drawing

AI summary

Systems, methods, and apparatus for using a multimodal response in the dynamic generation of client device output that is tailored to a current modality of a client device is disclosed herein. Multimodal client devices can engage in a variety of interactions across the multimodal spectrum including voice only interactions, voice forward interactions, multimodal interactions, visual forward interactions, visual only interactions etc. A multimodal response can include a core message to be rendered for all interaction types as well as one or more modality dependent components to provide a user with additional information.