Multimodal Action Orchestration With Unified Intent Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems process multi-modality inputs inefficiently, lacking seamless integration and interpretation of diverse user intents, and provide inflexible, static user interfaces that fail to adapt to user needs.

Innovation Solution

A unified schema is introduced to standardize user inputs and actions, leveraging LLMs and moderation rules for accurate intent mapping, and adaptive user interfaces that dynamically adjust based on action requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional systems handle different input modalities separately using specialized modules, then each modality can be processed with dedicated algorithms, but the processing pipeline becomes fragmented and inefficient

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidprocessing pipeline complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges multiple specialized modality processing modules into a single unified processing pipeline. The system receives inputs through various modalities (voice, image, text) and processes them through a common architecture that converts all inputs into a unified representation, eliminating the need for separate specialized modules for each modality and reducing overall system complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal processing pipeline that can handle multiple input modalities through a single standardized interface. The system uses a unified schema and common processing algorithms that work across all modalities, allowing the same processing infrastructure to serve multiple functions rather than requiring separate dedicated systems for each modality type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If traditional systems use predefined rules for mapping user intent to actions, then the system behavior is controlled and predictable, but the system becomes inflexible and fails to adapt to diverse user needs

Engineering Contradiction:
Improveintent mapping adaptabilityVSAvoidrule maintenance complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical system of predefined rules and hard-coded mappings with an AI-based intent recognition system. Instead of manually defining rules for intent-to-action mapping, the system uses machine learning models that automatically learn from data and adapt to user needs, eliminating the complexity of maintaining extensive rule sets while significantly improving adaptability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Ease of operation

If systems provide a one-size-fits-all static user interface, then the interface design is simple and consistent, but the interface fails to consider specific action requirements and user context

Engineering Contradiction:
Improveinterface usabilityVSAvoidinterface generation complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent transforms the static user interface into a dynamic, adaptive interface that automatically adjusts based on the triggered action and user context. The system generates customized interfaces for each action type, presenting only the relevant elements needed for that specific task rather than displaying a fixed comprehensive interface, thereby improving usability without requiring excessive complexity in the interface generation system.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12566736B2Orchestrating dynamic actions through multi-modality inputs
Publication Date: 2026.03.03 INTUIT INC
  • US12566736B2 patent drawing
  • US12566736B2 patent drawing
  • US12566736B2 patent drawing

AI summary

Certain aspects of the disclosure provide for processing multi-modality inputs to trigger dynamic actions. In examples, a method may include: receiving user input through one or more input modalities; converting the user input and one or more supported actions obtained from an actions repository into a unified schema; applying a set of moderation rules to the unified schema to generate a moderated input; generating a prompt based on the moderated input and one or more influencing examples; processing the prompt using one or more Large Language Models (LLMs) to obtain one or more matched actions and populated data models; augmenting the one or more matched actions with additional data; initiating one or more supported actions based on the one or more matched actions and populated data models; and generating a response to the user input based on results of executing the one or more supported actions.