Virtual Object Animation Generation via Linguistic Feature Sequences

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing end-to-end virtual object animation generation technologies are limited by reliance on specific audio features and lack generality, requiring manual labor and time for production, and cannot handle text input effectively, restricting user experience and production efficiency.

Innovation Solution

An end-to-end virtual object animation generation method that converts input information, including text or audio, into a pronunciation unit sequence, performs linguistic feature analysis, and uses a preset timing mapping model to generate animations independently of specific audio features, enabling automatic and general-purpose animation generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If end-to-end virtual object animation generation technology analyzes original input audio signal on acoustic basis, then automatic animation generation is achieved, but the technology relies on particular audio feature and is applicable only to dubber having particular speech feature, severely restricting generality

Engineering Contradiction:
Improveautomatic animation generationVSAvoidgenerality
Core Design Contradiction:
Extent of automationVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary processing layer that converts acoustic audio features into linguistic features through speech recognition and linguistic analysis. This intermediary step decouples the animation generation from direct acoustic dependency, allowing the system to generate animations based on linguistic content rather than acoustic characteristics, thereby achieving both automation and generality across different speakers

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the feature representation parameter from acoustic features to linguistic features. By transforming the input representation from raw audio waves to processed linguistic data (phonemes, words, sentences), the system eliminates dependence on particular audio features while maintaining automatic generation capability, thus resolving the contradiction between automation and adaptability

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If manual production process is used for virtual object animation, then high quality animation can be produced, but production process requires high labor and time costs

Engineering Contradiction:
Improveanimation qualityVSAvoidproduction efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent replaces the mechanical manual production process with an automated computational system. Instead of manual animation creation, the system uses speech recognition, linguistic analysis, and animation synthesis algorithms to automatically generate virtual object animations, thereby eliminating labor and time costs while maintaining quality through sophisticated processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Extent of automation

If only audio is used as input for virtual object animation generation, then automatic generation is achieved, but input selectivity is limited and user experience is affected

Engineering Contradiction:
Improveautomatic generationVSAvoidinput selectivity
Core Design Contradiction:
Extent of automationVSEase of operation

Solution Approach 1:

The patent makes the system universal by accepting multiple input types (audio, text, speech) through a unified processing architecture. The system can process audio signals, text inputs, and speech patterns through the same linguistic analysis pipeline, thereby enhancing input selectivity and user experience while maintaining automatic generation capability

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11810233B2End-to-end virtual object animation generation method and apparatus, storage medium, and terminal
Publication Date: 2023.11.07 MOFA (SHANGHAI) INFORMATION TECH CO LTD
  • US11810233B2 patent drawing
  • US11810233B2 patent drawing

AI summary

An end-to-end virtual object animation generation method includes receiving input information, where the input information includes text information or audio information of a virtual object animation to be generated; converting the input information into a pronunciation unit sequence; performing a feature analysis of the pronunciation unit sequence to obtain a corresponding linguistic feature sequence; and inputting the linguistic feature sequence into a preset timing mapping model to generate the virtual object animation based on the linguistic feature sequence.