Conversation-Driven Character Animation Using ML Persona Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional AI digital characters lack naturalness and nuance in interactions, failing to integrate verbal communications with non-verbal cues and respond dynamically to environmental and emotional factors, leading to static and unengaging simulations.

Innovation Solution

A system and method for producing conversation-driven character animation that uses machine learning models to predict the next state of a conversation, incorporating intent, sentiment, and interaction history, and generates animation streams that include verbal and non-verbal expressions, environmental conditions, and haptic effects in real-time, enabling dynamic and responsive interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional AI digital characters use a single synthesized persona, then the system is simple and easy to implement, but the interaction lacks naturalness and character variety

Engineering Contradiction:
Improveinteraction naturalnessVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments the digital character's persona into multiple distinct personas (e.g., cheerful, serious, playful) that can be dynamically selected and combined. Each persona represents a specific behavioral mode with associated non-verbal cues, allowing the character to adapt its personality to different conversation contexts while maintaining system manageability through modular persona design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from a static single persona to a dynamic multi-persona system where the character can switch between different personalities and emotional states in real-time. The persona selection is driven by conversation analysis, allowing the digital character to dynamically adapt its behavior, tone, and non-verbal expressions to match the interaction context and enhance naturalness.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If AI digital characters integrate multiple non-verbal cues and environmental factors, then the interaction becomes more natural and nuanced, but the computational complexity and processing requirements increase

Engineering Contradiction:
Improveconversation responsivenessVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system introduces an intermediary conversation analysis module that processes verbal communications and extracts intent, sentiment, and contextual information. This intermediary layer analyzes the conversation input and uses it to select appropriate non-verbal cues and persona configurations, thereby coordinating multiple factors (environmental conditions, emotional states, conversation content) without requiring direct complex interactions between all elements.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system employs a unified conversation analysis framework that simultaneously handles multiple functions: determining intent, analyzing sentiment, selecting personas, and choosing non-verbal cues. This multi-functional approach consolidates what would otherwise be separate processing systems into a single integrated module, reducing overall system complexity while maintaining comprehensive conversation responsiveness.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If the system generates real-time animation streams with multiple trained machine learning models, then the character interaction becomes more dynamic and engaging, but the computational resources and processing time increase

Engineering Contradiction:
Improveinteraction dynamismVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system pre-trains multiple machine learning models offline to recognize various conversation patterns, emotional states, and contextual scenarios. These pre-trained models are stored and ready for deployment, allowing the real-time system to quickly select and apply appropriate models based on the current conversation context without performing extensive training during interaction, thereby reducing processing time while maintaining high interaction dynamism.

Inventive Principle:
Principle #10Preliminary action

4Adaptability or versatility

If the digital character responds to environmental features and interaction goals, then the interaction becomes more context-aware and natural, but the system complexity and data processing requirements increase

Engineering Contradiction:
Improvecontext awarenessVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system uses a conversation analysis intermediary that serves as a mediator between environmental inputs and character responses. This intermediary module processes environmental features (location, weather, lighting), interaction goals, and conversation content, then translates them into persona selections and non-verbal cue configurations. This mediation simplifies the system architecture by providing a single point of integration for multiple context factors.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11983808B2Conversation-driven character animation
Publication Date: 2024.05.14 DISNEY ENTERPRISES INC
  • US11983808B2 patent drawing
  • US11983808B2 patent drawing
  • US11983808B2 patent drawing

AI summary

A system for producing conversation-driven character animation includes a computing platform having processing hardware and a system memory storing software code, the software code including multiple trained machine learning (ML) models. The processing hardware executes the software code to obtain a conversation understanding feature set describing a present state of a conversation between a digital character and a system user, and to generate an inference, using at least a first trained ML model of the multiple trained ML models and the conversation understanding feature set, the inference including labels describing a predicted next state of a scene within the conversation. The processing hardware further executes the software code to produce, using at least a second trained ML model of the multiple trained ML models and the labels, an animation stream of the digital character participating in the predicted next state of the scene within the conversation.