AI Video Generation Using Reference Motion for Frame Coherence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video generation methods using AI models to create images from text descriptions result in poor video coherence due to independent image generation processes, leading to significant jitter between frames.

Innovation Solution

A method that involves obtaining a content description text and a reference video, performing feature extraction to obtain semantic and action reference features, and generating a target video based on these features to enhance coherence.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If independent image generation processes are used to create video frames from text descriptions, then image generation flexibility is improved, but video coherence deteriorates due to significant jitter between consecutive frames

Engineering Contradiction:
Improveimage generation flexibilityVSAvoidvideo coherence
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The patent performs feature extraction from the reference video beforehand to obtain action reference features, which are then used to guide the generation of each video frame. This preliminary preparation of action templates ensures that each generated frame maintains consistency with the overall action sequence, thereby improving video coherence while preserving generation flexibility

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces a feedback mechanism where generated video frames are compared against the action reference features extracted from the reference video. This feedback loop ensures that each frame adheres to the overall action template, correcting deviations and maintaining temporal coherence across the video sequence

Inventive Principle:
Principle #23Feedback

2Stability of the object's composition

If action reference features from reference video are incorporated into the generation process, then video coherence is improved, but system complexity increases due to additional feature extraction and integration steps

Engineering Contradiction:
Improvevideo coherenceVSAvoidsystem complexity
Core Design Contradiction:
Stability of the object's compositionVSDevice complexity

Solution Approach 1:

The patent employs a unified video generation model that performs multiple functions: it processes text descriptions, integrates action reference features from reference videos, and generates coherent video frames. This multi-functional approach consolidates complexity into a single model rather than requiring separate modules for each function

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent transforms the reference video into action reference features through feature extraction, changing the parameter representation from raw video data to condensed action templates. This parameter transformation simplifies the integration process by working with compact feature representations rather than full video sequences

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250342616A1Video Generation Method and Apparatus, Storage Medium, and Electronic Device
Publication Date: 2025.11.06 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US20250342616A1 patent drawing
  • US20250342616A1 patent drawing
  • US20250342616A1 patent drawing

AI summary

This application discloses a video generation method and apparatus, a storage medium, and an electronic device. The method includes: obtaining a content description text and a content reference video, the content description text including information for describing target content expressed by a target video that is expected to be generated, and the content reference video including action reference information related to the target content; performing feature extraction on the content description text, to obtain text semantic features, the text semantic features being configured for representing semantic information of the content description text; performing feature extraction on the content reference video, to obtain video reference features, the video reference features being configured for representing the action reference information in the content reference video; and generating the target video based on the text semantic features and the video reference features. This application can improve quality of the generated target video.