Multimodal Response Generation for Dynamic Modality Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automated assistant systems struggle to dynamically generate client device output tailored to the current modality of a multimodal client device, leading to inefficiencies in resource usage and user interaction.

Innovation Solution

The system determines the current modality of a client device using sensor data and user interface input, and then selects appropriate components from a multimodal response to generate output, ensuring that output is tailored to the current modality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the system generates separate responses for each modality type, then each response can be optimized for its specific modality, but memory storage requirements and computational resources increase significantly

Engineering Contradiction:
Improvemodality-specific optimizationVSAvoidmemory storage requirements
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent merges multiple modality-specific responses into a single multimodal response structure that contains all possible output components (visual, audio, text, haptic) in one unified data structure. This single response is then dynamically adapted at runtime based on the current device modality, eliminating the need to store multiple separate response files while maintaining the ability to optimize for each modality.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The multimodal response structure serves multiple functions simultaneously: it acts as a template for all modality types, a runtime configuration source for dynamic adaptation, and a storage-efficient single-source data structure. This universal structure can generate optimized output for any modality type without requiring separate specialized responses for each.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If the system processes and selects from multiple modality-specific responses, then output can be tailored to current modality, but computational efficiency decreases and processing time increases

Engineering Contradiction:
Improvedynamic modality adaptationVSAvoidcomputational efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system performs preliminary organization of all modality-specific output components within a single multimodal response structure during response generation. The modality-specific components are pre-labeled and structured in advance, so that at runtime the system only needs to perform simple selection and filtering operations rather than complex processing of multiple separate responses.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If the system transmits separate responses for each modality from remote server, then each response can be optimized, but network bandwidth usage and transmission time increase

Engineering Contradiction:
Improvemodality-optimized contentVSAvoidtransmission time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent combines all modality-specific response content into a single multimodal response that is transmitted over the network. This single transmission contains all necessary components (visual, audio, text, haptic) organized by modality type, eliminating the need for multiple separate network transmissions while still allowing the client device to select and render only the components appropriate for the current modality.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12327559B2Multimodal responses
Publication Date: 2025.06.10 GOOGLE LLC
  • US12327559B2 patent drawing
  • US12327559B2 patent drawing
  • US12327559B2 patent drawing

AI summary

Systems, methods, and apparatus for using a multimodal response in the dynamic generation of client device output that is tailored to a current modality of a client device is disclosed herein. Multimodal client devices can engage in a variety of interactions across the multimodal spectrum including voice only interactions, voice forward interactions, multimodal interactions, visual forward interactions, visual only interactions etc. A multimodal response can include a core message to be rendered for all interaction types as well as one or more modality dependent components to provide a user with additional information.