Avatar Animation Using Text and Neural Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing avatar creation methods are limited by requiring specific video inputs, leading to incomplete or inaccurate representations of users, especially when video inputs are not available, and fail to convey certain human emotions effectively.

Innovation Solution

A method using a neural network to generate speech data sets and movement parameters for avatars based on received text and determined emotional states, allowing for more accurate and flexible avatar creation, including in situations where only textual input is available.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If video input is required for avatar creation, then avatar accuracy is improved, but system flexibility and ease of operation deteriorate

Engineering Contradiction:
Improveavatar accuracyVSAvoidinput method flexibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces text as an intermediary input modality that bridges the gap between limited video input availability and the need for accurate avatar creation. When video input is unavailable or insufficient, text descriptions serve as a mediator to convey user characteristics, enabling avatar generation without direct video capture.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system dynamically changes the input parameters accepted by the avatar creation system. Instead of requiring fixed video input parameters, the system accepts variable parameters including text descriptions, emotional state data, and contextual information, allowing flexible adaptation to different input scenarios while maintaining avatar generation capability.

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If video input with specific features is required, then avatar completeness is improved, but ease of operation and accessibility deteriorate

Engineering Contradiction:
Improveavatar completenessVSAvoidinput accessibility
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

Instead of requiring users to provide complex video inputs with specific features, the system inverts the approach by accepting simple text descriptions and generating comprehensive avatar data. This inversion simplifies the user's task while maintaining or improving the completeness of the resulting avatar through advanced processing of the text input.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The system performs self-service by automatically extracting and inferring avatar characteristics from text input without requiring users to manually provide detailed video footage. The neural network processes the text and automatically generates movement parameters, emotional states, and avatar representations, reducing the operational burden on users.

Inventive Principle:
Principle #25Self-service

3Device complexity

If conventional motion ranges are used for avatar animation, then device complexity is reduced, but emotional expression accuracy deteriorates

Engineering Contradiction:
Improveanimation system complexityVSAvoidemotional expression accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent implements dynamic motion ranges that adapt to the detected emotional state. Instead of using fixed conventional motion ranges, the system adjusts the amplitude, speed, and type of avatar movements based on the emotional context derived from text input, enabling more accurate emotional expression while maintaining manageable system complexity through rule-based adjustments.

Inventive Principle:
Principle #15Dynamics

4Adaptability or versatility

If text input is accepted instead of video, then system flexibility is improved, but manufacturing precision of avatar representation deteriorates

Engineering Contradiction:
Improveinput method flexibilityVSAvoidavatar representation accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent substitutes the mechanical video capture and processing system with a neural network-based text processing system. This replacement uses artificial intelligence to interpret text descriptions and generate accurate avatar representations, achieving precision through computational methods rather than direct video measurement.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11593984B2Using text for avatar animation
Publication Date: 2023.02.28 APPLE INC
  • US11593984B2 patent drawing
  • US11593984B2 patent drawing
  • US11593984B2 patent drawing

AI summary

Systems and processes for animating an avatar are provided. An example process of animating an avatar includes at an electronic device having one or more processors and memory, receiving text, determining an emotional state, and generating, using a neural network, a speech data set representing the received text and a set of parameters representing one or more movements of an avatar based on the received text and the determined emotional state.