Video Generation Method for Dynamic Character Emotion Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current image driving technologies struggle to effectively generate dynamic character videos from static images based on speech segments, lacking flexibility and realism in character emotion expression.
Innovation Solution
A method and apparatus that change the character emotion of a static image using a character emotion feature to obtain a target image, which is then driven by a character driving network based on a speech segment to generate a video, employing an emotion editing network and a face driving network to enhance realism and flexibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If current image driving technologies are used to generate dynamic character videos from static images, then the basic video generation function is achieved, but the flexibility and realism in character emotion expression deteriorates
Solution Approach 1:
The system segments the emotion expression process into two independent stages: first, emotion editing of the static character image to generate an edited character image with desired emotion; second, driving the edited image to generate video. This segmentation allows independent optimization of each stage, improving both realism in emotion expression and flexibility in controlling different emotion types.
Solution Approach 2:
The system introduces dynamic emotion control by allowing users to specify different emotion types for video generation. The emotion editing network dynamically adjusts the static character image's emotion based on input emotion types, and the driving network dynamically generates video frames that reflect these emotion changes, achieving both realism and flexibility.
2Adaptability or versatility
If a static character image is directly driven by speech segment, then the video generation process is simple, but the character emotion expression lacks variety and personalization
Solution Approach 1:
The system performs preliminary emotion editing on the static character image before the driving process. By pre-adjusting the emotion of the character image according to the desired emotion type, the subsequent video generation can focus on motion and expression dynamics, achieving emotion variety without significantly increasing overall process complexity.
3Reliability
If emotion editing is performed on the original character image, then the target character image with desired emotion is obtained, but the processing time and computational resources increase
Solution Approach 1:
The system extracts and isolates the emotion editing function into a separate network module. This dedicated emotion editing network focuses specifically on adjusting emotion attributes of the character image, improving the accuracy of emotion expression while allowing the main driving network to operate efficiently on the pre-edited image, thereby balancing accuracy and processing time.
Data Source
AI summary
Provided are a video generation method and apparatus, a device and a storage medium, relating to the field of artificial intelligence and, in particular, to the fields of computer vision and deep learning. The method includes changing a character emotion of an original character image according to a character emotion feature of a to-be-generated video to obtain a target character image; and driving the target character image by use of a character driving network and based on a speech segment to obtain the to-be-generated video.


