Speech-Driven Lip Video Generation With Smooth Frame Transitions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video generation methods using speech-driven lip shapes result in abrupt transitions between frames, leading to poor lip shape driving effects in generated videos.

Innovation Solution

A video generation method that involves acquiring target audio and video data, performing mask processing, feature extraction using a target multimodal model, and predicting lip areas based on synchronized audio and video features to ensure natural transitions in lip shape across frames.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If image generation model is used to generate lip shape from target speech, then lip shape can be generated, but the transition of lip shape among different video frames is abrupt

Engineering Contradiction:
Improvelip shape generation accuracyVSAvoidlip shape transition smoothness
Core Design Contradiction:
Manufacturing precisionVSStability of the object's composition

Solution Approach 1:

The patent segments the lip shape generation process into multiple components: obtaining target audio data, obtaining first video data, performing mask processing to obtain second video data, extracting features from both video data sources, and finally predicting lip areas. This segmentation allows each component to be optimized independently, ensuring both accuracy and temporal smoothness.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary mechanism by using both first video data and second video data (with mask processing) as intermediate representations. These intermediaries capture different aspects of lip movement and are combined through feature extraction to produce smooth transitions while maintaining generation accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If mask processing is performed on lip area in video data, then lip area can be isolated for processing, but additional processing steps are required

Engineering Contradiction:
Improvelip area identification accuracyVSAvoidvideo generation process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies mask processing as a preliminary action to obtain second video data with isolated lip areas before the main prediction process. This preliminary segmentation improves the precision of lip area identification and enables more accurate feature extraction, justifying the additional processing step.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If synchronization alignment training is performed on paired sample audio and sample video, then audio-video synchronization can be improved, but training complexity increases

Engineering Contradiction:
Improveaudio-video synchronization reliabilityVSAvoidmodel training complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent performs synchronization alignment training as a preliminary action during the model training phase, before actual video generation. This preliminary training ensures that the target multimodal model learns to align audio and video features effectively, improving synchronization reliability in the generated videos while the actual generation process remains relatively simple.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250392796A1Video generation method, apparatus, device, medium and program product
Publication Date: 2025.12.25 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US20250392796A1 patent drawing
  • US20250392796A1 patent drawing
  • US20250392796A1 patent drawing

AI summary

The present disclosure relates to the technical field of video processing, and discloses a video generation method, apparatus, device, medium and program product. The method includes: acquiring target audio data and first video data of a target object; acquiring second video data, the second video data is obtained by performing mask processing on a lip area in video data of the target object; performing feature processing on the target audio data based on a target multimodal model to obtain a target audio feature; performing feature extraction on the first video data and the second video data to obtain a feature to be processed; and predicting a lip area in the second video data based on the target audio feature and the feature to be processed, to determine a target video corresponding to the target audio data.