Virtual Role Multimodal Interaction for Complex Situations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current intelligent assistants suffer from low information acquisition efficiency and poor interaction experience due to limited input modalities and rigid output formats, primarily relying on text and speech, which fail to handle complex situations effectively.

Innovation Solution

A virtual role-based multimodal interaction method that processes input information of various types through a perception layer and logic decision-making layer to generate multimodal virtual content, including virtual roles, scenes, and effects, enhancing interaction through speech, actions, and diverse outputs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If existing intelligent assistants use only text and speech interaction modalities, then the system complexity is low, but the information acquisition efficiency is low and interaction experience is poor

Engineering Contradiction:
Improveinformation acquisition efficiencyVSAvoidinput modality complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The intelligent assistant system is enhanced to support multiple input modalities including text, speech, images, and audio, allowing a single system to perform diverse interaction functions. The perception layer incorporates multiple recognition modules (speech recognition, text recognition, image recognition) that enable the system to process different types of input data through unified processing pipelines, thereby improving information acquisition efficiency without requiring separate systems for each modality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system architecture is segmented into distinct functional layers: perception layer for multi-modal input processing, logic decision-making layer for processing recognition results, and output layer for generating responses. This segmentation allows each layer to specialize in specific tasks, improving overall processing efficiency while maintaining manageable system complexity through modular design.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If existing intelligent assistants use only text and speech output formats, then the device complexity is low, but the interaction experience is poor and presentation is rigid

Engineering Contradiction:
Improveoutput format versatilityVSAvoidoutput modality complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The output layer is designed to support multiple output modalities including text, speech, images, and video, enabling the intelligent assistant to adapt its response format based on the input type and user needs. This multi-functional output capability enhances interaction experience by providing diverse presentation formats while maintaining a unified system architecture that manages complexity through standardized processing interfaces.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If the dialogue system module processes only simple question answering, then the processing speed is fast, but the system cannot handle complex situations and gives irrelevant answers

Engineering Contradiction:
Improvecomplex situation handling capabilityVSAvoidlogic decision-making complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The logic decision-making layer is segmented into specialized processing modules that handle different types of tasks. The system can route simple question-answering queries through streamlined processing paths while directing complex situations to more comprehensive analysis modules. This segmentation enables the system to adapt its processing depth to the complexity of the input, improving both handling capability and efficiency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts its processing behavior based on the complexity and type of input received. For simple queries, the system uses efficient direct processing; for complex situations involving multiple data types or requiring cross-modal analysis, the system activates more comprehensive processing routines. This dynamic adaptation allows the system to handle varying levels of complexity without requiring maximum processing capacity for all inputs.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12367640B2Virtual role-based multimodal interaction method, apparatus and system, storage medium, and terminal
Publication Date: 2025.07.22 MOFA (SHANGHAI) INFORMATION TECH CO LTD
  • US12367640B2 patent drawing
  • US12367640B2 patent drawing
  • US12367640B2 patent drawing

AI summary

A virtual role-based multimodal interaction method, apparatus and system, a storage medium, and a terminal, where the method includes: acquiring input information, where the input information includes one or more data types; inputting the input information into a perception layer to enable the perception layer to recognize and process the input information according to a data type of the input information to obtain a recognition result; inputting the recognition result into a logic decision-making layer to enable the logic decision-making layer to process the recognition result and generate a drive instruction corresponding to the recognition result; acquiring multimodal virtual content according to the drive instruction, wherein the multimodal virtual content includes at least a virtual role; and outputting the acquired multimodal virtual content.