Diffusion Video Generation Guided by Keyframes and Text

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video generation methods using machine learning models often produce videos with poor dynamicity and lack realistic motion visual effects, struggling to create high-dynamic action videos and complex camera movements.

Innovation Solution

A video generation method utilizing a diffusion model that combines image instructions of the first and last frames with text instructions to generate videos, enhancing dynamic visual effects by specifying multiple images and text for a target video.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If a machine learning model is used to generate videos from a pre-specified image and text description, then video generation can be automated, but the generated videos exhibit poor dynamicity with objects lacking obvious actions and dynamic effects

Engineering Contradiction:
Improvevideo generation automationVSAvoiddynamic visual effect quality
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent segments the video generation process into multiple independent modules: an image processing module that extracts features from input images, a text processing module that encodes text descriptions, and a video generation module that synthesizes the final video. This segmentation allows each module to specialize in specific tasks, improving overall video quality and dynamicity while maintaining automation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate representations as mediators between input data and final video output. Specifically, it uses image features extracted from input images and text embeddings from text descriptions as intermediate representations that guide the video generation process, enabling more controlled and dynamic video synthesis.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of manufacture

If existing video generation methods are used, then video content can be generated from text descriptions, but the videos lack realistic motion visual effects and complex camera movements

Engineering Contradiction:
Improvevideo generation convenienceVSAvoidmotion visual effect realism
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent changes key parameters in the video generation process by using learned image features and text embeddings instead of direct pixel manipulation. It also employs a diffusion model with adjustable noise schedules and generation steps, allowing fine-tuned control over motion dynamics and visual realism while maintaining ease of use through text descriptions.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If simple image-text pairs are used for video generation, then the generation process is simple, but the videos cannot achieve complex scenes and movements

Engineering Contradiction:
Improvegeneration process complexityVSAvoidvideo content complexity
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent transitions from 2D image processing to 3D video generation by introducing temporal dimension. It processes images and text in the spatial domain while generating videos that evolve in the temporal domain, enabling complex scenes and movements through time-based transformations guided by the input image and text description.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12562194B2Method, apparatus, device and medium for generating a video
Publication Date: 2026.02.24 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US12562194B2 patent drawing
  • US12562194B2 patent drawing
  • US12562194B2 patent drawing

AI summary

Provided are a method, apparatus, device and medium for generating a video. In one method, a plurality of images for respectively describing a plurality of target images in a target video are received. A text for describing a content of the target video is received. The target video is generated based on the plurality of images and the text according to a generation model. With exemplary implementations of the present disclosure, the plurality of images received can serve as guiding data to determine a development direction of a story in the video, which contributes to the generation of a richer and more realistic dynamic video.