Multimodal Input Processing for Virtual Agents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing virtual agents and AI-powered systems are limited in handling complex, natural language interactions and lack the ability to effectively process multimodal inputs such as speech, sensor data, or visual inputs, leading to less engaging and less effective user-agent interactions.

Innovation Solution

A method and system for multimodal input processing using a Generative AI model that identifies principal entities within the input, extracts information, and generates responses, while also adapting to user preferences and emotions, and continuously improving over time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If text-based rule-based systems are used for virtual agents, then implementation is simple and maintenance is easy, but the ability to handle complex natural language interactions is limited

Engineering Contradiction:
Improveability to handle complex natural language interactionsVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent replaces traditional rule-based mechanical systems with neural network-based AI models that can automatically learn and adapt to complex language patterns, enabling the virtual agent to handle nuanced natural language interactions without manual rule configuration

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system dynamically adjusts processing parameters such as attention weights, temperature, and top-k sampling based on the complexity of user input, allowing the model to adapt its behavior to match the required level of response sophistication

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If traditional text-based processing is used, then processing speed is fast and resource consumption is low, but the ability to process multimodal inputs is limited

Engineering Contradiction:
Improveability to process multimodal inputsVSAvoidcomputational resource consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent divides multimodal input processing into separate specialized modules for text, speech, images, and sensor data, each processed by dedicated sub-models before being integrated, allowing efficient resource utilization by only activating necessary processing pipelines

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system employs a unified transformer architecture that can process multiple input modalities through the same core processing mechanisms, enabling the model to handle diverse input types without requiring entirely separate processing systems

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If basic NLP models are used, then response generation is quick, but understanding of user emotions and context is insufficient

Engineering Contradiction:
Improveunderstanding accuracyVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary processing of user input including emotion detection, intent classification, and entity extraction before main response generation, preparing contextual information in advance to accelerate the subsequent response synthesis process

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces intermediate representation layers that transform raw multimodal inputs into standardized contextual embeddings, serving as a bridge between diverse input formats and the response generation mechanism, improving both accuracy and processing efficiency

Inventive Principle:
Principle #24Intermediary (Mediator)

4Adaptability or versatility

If static response patterns are used, then system simplicity is maintained, but user engagement and satisfaction decrease

Engineering Contradiction:
Improvecommunication adaptabilityVSAvoidmodel complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system dynamically adapts its communication style, tone, and level of formality based on user preferences, emotional state, and interaction history, transforming from static response patterns to dynamic adaptive communication that evolves with each interaction

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent implements feedback loops where user responses are analyzed to refine future interactions, with the model learning from engagement metrics and adjusting its communication strategy to improve user satisfaction over time

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250165744A1Method and system for integrated multimodal input processing for virtual agents
Publication Date: 2025.05.22 QUANTIPHI INC
  • US20250165744A1 patent drawing
  • US20250165744A1 patent drawing
  • US20250165744A1 patent drawing

AI summary

A method and system for multimodal input processing for a virtual agent is provided herein. The method comprises obtaining a multimodal input by the virtual agent from a user. The method further comprises identifying a plurality of principal entities within the multimodal input. The method further comprises extracting information about each entity of the plurality of principal entities. Further, the method comprises generating a response based on the extracted information.