Face Video Generation Using Segmented Mouth Shape Drivers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current human face and mouth shape driving solutions generate videos in a generic style, failing to accurately reflect the personal mouth shape style of the target object, resulting in low accuracy of the generated human face video.

Innovation Solution

A method involving the use of a mouth-shape multimedia resource and a reference human face image to extract mouth-shape driving features, generate stylistic human face images, and create a stylistic human face video that reflects the personal mouth shape style of the target object, utilizing a trained human face and mouth shape driver model with a feature extraction network and a human face driving network.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a general-purpose human face and mouth shape driver model is used, then the device complexity is reduced, but the manufacturing precision of the generated video deteriorates

Engineering Contradiction:
Improvemodel complexityVSAvoidvideo accuracy
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The driver model is segmented into multiple specialized sub-models: a mouth shape driver model for lip movements, an eye driver model for eye movements, and a facial expression driver model for facial expressions. Each sub-model specializes in a specific aspect, improving overall video accuracy without requiring a single overly complex general-purpose model.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different parts of the face are processed with different levels of detail and specialized models. The mouth region uses a dedicated mouth shape driver model, eyes use an eye driver model, and other facial regions use the facial expression driver model. This local specialization ensures high precision for each facial component while maintaining manageable overall system complexity.

Inventive Principle:
Principle #3Local quality

2Ease of operation

If a general-purpose human face and mouth shape driver model is used, then the ease of operation is improved, but the reliability of the generated video deteriorates

Engineering Contradiction:
Improveoperation simplicityVSAvoidvideo authenticity
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The driver model is divided into specialized sub-models for different facial components (mouth, eyes, expressions), where each sub-model is trained on specific data and handles particular aspects of facial animation. This segmentation improves reliability by ensuring each component is driven by a model optimized for its specific function, while the integrated system remains easy to operate through unified input processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system employs multiple loss functions during training that provide feedback on different aspects of video quality: content loss for structural accuracy, perceptual loss for visual realism, and style loss for capturing individual characteristics. This multi-faceted feedback mechanism ensures high reliability and authenticity in the generated videos while maintaining ease of operation.

Inventive Principle:
Principle #23Feedback

3Manufacturing precision

If feature extraction and style transfer processes are implemented, then the manufacturing precision of the video is improved, but the productivity of the generation process deteriorates

Engineering Contradiction:
Improvevideo qualityVSAvoidgeneration speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

Style vectors are extracted and stored in advance from reference images of target objects. During video generation, these pre-extracted style vectors are directly applied without requiring real-time style transfer computations. This preliminary action significantly improves generation speed while maintaining high video quality through the use of pre-computed stylistic features.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The style transfer process is separated into a distinct preprocessing step where style vectors are extracted from reference images before the main video generation process. This extraction is performed once and reused across multiple video generations, reducing the computational burden during actual video synthesis and improving overall productivity while maintaining manufacturing precision.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20240420403A1Face video generation method, device and electronic equipment
Publication Date: 2024.12.19 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US20240420403A1 patent drawing
  • US20240420403A1 patent drawing
  • US20240420403A1 patent drawing

AI summary

A method for generating a human face video, an apparatus for generating a human face video, includes: obtaining a mouth-shape multimedia resource and a reference human face image of a target object; obtaining a reference style vector of the target object; for each resource frame in the mouth-shape multimedia resource, obtaining a mouth-shape driving feature by performing a feature extraction process on the resource frame; generating a stylistic human face image corresponding to the resource frame according to the mouth-shape driving feature, the reference human face image, and the reference style vector; and determining a stylistic human face video of the target object.