Video Generation Method for Dynamic Character Emotion Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current image driving technologies struggle to effectively generate dynamic character videos from static images based on speech segments, lacking flexibility and realism in character emotion expression.

Innovation Solution

A method and apparatus that change the character emotion of a static image using a character emotion feature to obtain a target image, which is then driven by a character driving network based on a speech segment to generate a video, employing an emotion editing network and a face driving network to enhance realism and flexibility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If current image driving technologies are used to generate dynamic character videos from static images, then the basic video generation function is achieved, but the flexibility and realism in character emotion expression deteriorates

Engineering Contradiction:
Improverealism in character emotion expressionVSAvoidflexibility in character emotion expression
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system segments the emotion expression process into two independent stages: first, emotion editing of the static character image to generate an edited character image with desired emotion; second, driving the edited image to generate video. This segmentation allows independent optimization of each stage, improving both realism in emotion expression and flexibility in controlling different emotion types.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces dynamic emotion control by allowing users to specify different emotion types for video generation. The emotion editing network dynamically adjusts the static character image's emotion based on input emotion types, and the driving network dynamically generates video frames that reflect these emotion changes, achieving both realism and flexibility.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If a static character image is directly driven by speech segment, then the video generation process is simple, but the character emotion expression lacks variety and personalization

Engineering Contradiction:
Improvevariety in character emotion expressionVSAvoidcomplexity of video generation process
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs preliminary emotion editing on the static character image before the driving process. By pre-adjusting the emotion of the character image according to the desired emotion type, the subsequent video generation can focus on motion and expression dynamics, achieving emotion variety without significantly increasing overall process complexity.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If emotion editing is performed on the original character image, then the target character image with desired emotion is obtained, but the processing time and computational resources increase

Engineering Contradiction:
Improveaccuracy of character emotion expressionVSAvoidvideo generation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system extracts and isolates the emotion editing function into a separate network module. This dedicated emotion editing network focuses specifically on adjusting emotion attributes of the character image, improving the accuracy of emotion expression while allowing the main driving network to operate efficiently on the pre-edited image, thereby balancing accuracy and processing time.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11836837B2Video generation method, device and storage medium
Publication Date: 2023.12.05 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11836837B2 patent drawing
  • US11836837B2 patent drawing
  • US11836837B2 patent drawing

AI summary

Provided are a video generation method and apparatus, a device and a storage medium, relating to the field of artificial intelligence and, in particular, to the fields of computer vision and deep learning. The method includes changing a character emotion of an original character image according to a character emotion feature of a to-be-generated video to obtain a target character image; and driving the target character image by use of a character driving network and based on a speech segment to obtain the to-be-generated video.