AI Animation Character Drive System for Immersive Interaction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current human-computer interaction methods lack the immersive experience for users, as they primarily rely on text or speech inputs without sophisticated animation character drive methods to simulate realistic expressions and sounds.
Innovation Solution
An AI-based animation character drive method that determines a first expression base from media data of a speaker's facial expressions, then uses this base to drive a second animation character by simulating the corresponding sounds and expressions based on target text information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If traditional speech interaction methods are used, then the system is simple to operate, but the user immersion and sense of presence are insufficient
Solution Approach 1:
The patent creates a virtual copy of the speaker's facial expressions and speech characteristics by extracting features from media data and reproducing them through an animation character. This allows the animation character to mimic the speaker's expressions, mouth shapes, and acoustic features, thereby enhancing user immersion while maintaining simple speech-based interaction.
Solution Approach 2:
The patent introduces an animation character as an intermediary between the user and the system response. This intermediary enriches the interaction by displaying facial expressions and speech patterns, bridging the gap between simple text/speech input and immersive communication experience.
2Reliability
If sophisticated animation character drive methods are implemented, then the sense of immersion is improved, but the device complexity increases
Solution Approach 1:
The patent divides the complex animation character drive system into separate functional modules: feature extraction from media data, expression base determination, acoustic feature analysis, and animation character rendering. This segmentation allows each module to handle specific tasks independently, managing overall system complexity while achieving sophisticated immersion effects.
Solution Approach 2:
The patent transforms complex media data into standardized parameters including expression parameters, acoustic features, and mouth shape parameters. By converting raw data into structured parameters that can be directly applied to animation characters, the system achieves realistic expressions without requiring overly complex processing pipelines.
3Manufacturing precision
If expression bases from media data are used, then the facial expression realism is improved, but the processing time and computational resources increase
Solution Approach 1:
The patent performs preliminary extraction of expression bases and acoustic features from media data in advance, creating reusable expression templates and parameter sets. This pre-processing allows the system to quickly apply pre-defined expression patterns during real-time interaction, reducing processing time while maintaining high expression realism.
Solution Approach 2:
The patent uses lightweight expression parameters and acoustic features that can be quickly processed and discarded after use, rather than maintaining complex persistent models. This approach reduces computational overhead and processing time while still achieving realistic facial expressions for each interaction instance.
Data Source
Figure 1
Figure 2~3
Figure 4~5
AI summary
An animation image driving method based on artificial intelligence, and a related device. The method comprises: collecting media data of facial expression changes when a speaker speaks a speech, and determining a first expression base of a first animation image corresponding to the speaker, wherein the first expression base can reflect different expressions of the first animation image; after target text information used for driving a second animation image is determined, determining an acoustic feature and a target expression parameter corresponding to the target text information according to the target text information, the collected media data and the first expression base; and driving the second animation image having a second expression base by means of the acoustic feature and the target expression parameter, so that the second animation image can give out the sound of the target text information spoken by the speaker by means of acoustic feature simulation, and a facial expression conforming to the due expression of the speaker is made in a sounding process, so that vivid substitution feeling and immersion feeling are brought to a user, and the interaction experience of the user and the animation image is improved.