Virtual Character Video Generation with Emotional Response

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for generating virtual characters fail to create realistic interactions by ignoring emotional responses and facial expressions that are commensurate with the speech spoken by speakers, resulting in virtual characters that lack the verisimilitude of human communication.

Innovation Solution

A computer-implemented method using machine learning techniques to generate videos of virtual characters that exhibit emotional reactions and blinking motions responsive to speech, facial expressions, and head pose motions of a speaker, by processing audio and visual features through machine learning models to create a discrete latent space and sequence of blink coefficients.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional methods are used to generate virtual characters, then the generation process is simple, but the virtual characters lack realistic emotional responses and facial expressions

Engineering Contradiction:
Improverealism of virtual character interactionsVSAvoidcomplexity of processing system
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system segments the processing into distinct modules: audio feature extraction, visual feature extraction, emotion recognition, and video generation. Each module handles specific tasks independently, allowing complex processing to be broken down into manageable components that can be optimized separately.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer that includes emotion recognition models and feature extraction modules. These intermediaries bridge the gap between raw input data (audio and visual) and the final virtual character output, transforming and interpreting the data before generation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If emotional responses and facial expressions are added to virtual characters, then the verisimilitude of human communication is improved, but the processing time and computational resources increase

Engineering Contradiction:
Improveverisimilitude of human communicationVSAvoidprocessing time for video generation
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary processing by extracting audio features, visual features, and emotion data before the main video generation process. This pre-computation of intermediate representations allows the actual video synthesis to be faster, as it works with prepared data rather than raw inputs.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent employs dynamic processing where the system adjusts the level of detail and processing intensity based on the complexity of the input. The emotion recognition and facial expression generation are dynamically activated only when needed, optimizing the balance between realism and processing time.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If the system processes audio and visual features through multiple machine learning models, then the accuracy of emotional reaction generation is improved, but the system complexity increases

Engineering Contradiction:
Improveaccuracy of emotional reaction detectionVSAvoidnumber of machine learning models
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent designs a multi-functional processing architecture where a single video generation model can handle multiple tasks: generating facial expressions, eye movements, head poses, and emotional reactions. This multi-functionality reduces the need for separate specialized models for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses parameter changes in machine learning models to adapt to different input conditions. By adjusting model parameters and processing thresholds dynamically, the system maintains high accuracy across diverse scenarios without requiring separate models for each specific task or demographic.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240346735A1System and method for generating videos depicting virtual characters
Publication Date: 2024.10.17 UNIVERSITY OF ROCHESTER
  • US20240346735A1 patent drawing
  • US20240346735A1 patent drawing
  • US20240346735A1 patent drawing

AI summary

Features described herein pertain to generative machine learning, and more particularly, to machine learning techniques for generating virtual characters. A video that depicts a first subject and includes an audio component that corresponds to speech spoken by the first subject and an image that depicts a second subject are provided to and used by one or more machine learning models to generate a video that depicts the second subject. The second subject can blink and exhibit emotional characteristic and reactions that are responsive to the speech spoken by the first subject and/or a characteristic of the first subject such as a facial expression and/or head pose motion. The generated video can be displayed and/or stored where it can be later retrieved.