Multi-modal Model for Virtual Character Interaction Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing virtual character systems often misinterpret user inputs due to isolated processing of voice and facial data, leading to inaccurate responses and a decreased user experience.

Innovation Solution

A multi-modal model that combines speech recognition, natural language understanding, facial expression recognition, environmental awareness, and knowledge bases to dynamically generate accurate animations and actions for virtual characters, using error correction techniques to validate user inputs and improve interaction accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If isolated processing of voice and facial data is used, then device complexity is reduced, but measurement precision deteriorates

Engineering Contradiction:
Improveprocessing complexityVSAvoidinput interpretation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent combines multiple isolated processing systems (voice recognition, facial expression recognition, environmental awareness) into a unified multi-modal processing framework. This integration allows the virtual character to simultaneously analyze speech, facial expressions, and environmental context, thereby improving input interpretation accuracy while managing complexity through systematic architecture.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The processing system is designed to handle multiple types of input data (audio, visual, environmental) through a unified multi-modal model. This multi-functional approach enables the same system to process diverse data types concurrently, improving overall measurement precision without requiring separate dedicated systems for each input modality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If multiple internal models are implemented, then response accuracy is improved, but device complexity increases

Engineering Contradiction:
Improveresponse accuracyVSAvoidmodel integration complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent divides the complex processing task into distinct internal models, each specialized for a specific function: speech recognition model, natural language understanding model, facial expression recognition model, and environmental awareness model. This segmentation allows each model to focus on its specific task, improving overall response accuracy while making the system more manageable through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The multi-modal model acts as an intermediary that receives inputs from multiple internal models, integrates their outputs, and generates the final virtual character response. This intermediary layer coordinates the interactions between different models, managing complexity by providing a unified interface for integrating diverse model outputs.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If error correction techniques are used, then interaction accuracy is improved, but processing time increases

Engineering Contradiction:
Improveuser input accuracyVSAvoidprocessing delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary validation of user inputs using multiple internal models before final processing. By pre-checking inputs through speech recognition, facial expression analysis, and environmental context verification, the system identifies and corrects potential errors early in the processing pipeline, improving overall input accuracy while minimizing processing delays through efficient early validation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The multi-modal processing system incorporates feedback loops where the output from one internal model influences the processing in other models. For example, facial expression recognition feedback can adjust speech interpretation, and environmental context feedback can refine natural language understanding. This feedback mechanism improves input accuracy by cross-validating interpretations across different modalities.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11501480B2Multi-modal model for dynamically responsive virtual characters
Publication Date: 2022.11.15 LONDON ALLEY ENTERTAINMENT LLC
  • US11501480B2 patent drawing
  • US11501480B2 patent drawing
  • US11501480B2 patent drawing

AI summary

The disclosed embodiments relate to a method for controlling a virtual character (or “avatar”) using a multi-modal model. The multi-modal model may process various input information relating to a user and process the input information using multiple internal models. The multi-modal model may combine the internal models to make believable and emotionally engaging responses by the virtual character. The link to a virtual character may be embedded on a web browser and the avatar may be dynamically generated based on a selection to interact with the virtual character by a user. A report may be generated for a client, the report providing insights as to characteristics of users interacting with a virtual character associated with the client.