Video Diffusion Model for Full-Duration Generation and Motion Coherence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video generation methods, particularly text-to-video (T2V) and spatial super-resolution (SSR) models, face high computational costs and memory consumption due to the high dimensionality of video data, leading to limited global coherence and motion inconsistencies.

Innovation Solution

A machine-learned denoising diffusion model performs temporal downsampling and upsampling operations to generate multiple frames simultaneously, combined with a spatial super-resolution model applied over smaller temporal windows, optimizing computational efficiency and maintaining global motion coherence.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If existing T2V models are used to generate video frames, then video generation is achieved, but computational costs and memory consumption are excessively high

Engineering Contradiction:
Improvevideo generation capabilityVSAvoidcomputational cost
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent segments the video generation process into two distinct stages: (1) a denoising diffusion model generates low-resolution video frames at reduced computational cost, and (2) a spatial super-resolution model enhances selected key frames to high resolution. This segmentation allows the computationally intensive high-resolution generation to be applied selectively rather than to all frames, significantly reducing overall computational requirements while maintaining visual quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by differentiating the resolution requirements across different temporal regions of the video. Instead of uniformly generating all frames at high resolution, the system identifies and applies super-resolution enhancement only to key frames that require high visual fidelity, while intermediate frames remain at lower resolution. This localized application of high quality processing optimizes the balance between computational cost and perceived video quality.

Inventive Principle:
Principle #3Local quality

2Productivity

If existing T2V models are used to generate video frames, then video generation is achieved, but global motion coherence is limited due to temporal aliasing ambiguities

Engineering Contradiction:
Improvevideo generation capabilityVSAvoidglobal motion coherence
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies preliminary action by having the denoising diffusion model generate low-resolution video frames first, establishing the global motion coherence and temporal consistency across the entire video sequence. Once the temporal structure is firmly established at low resolution, the spatial super-resolution model then enhances key frames without disrupting the previously established motion coherence, thereby maintaining global consistency while improving local visual quality.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If spatial super-resolution models are applied to all video frames simultaneously, then high-resolution output is achieved, but memory consumption becomes substantial

Engineering Contradiction:
Improvevideo resolution qualityVSAvoidmemory consumption
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent segments the set of video frames into different categories: key frames that require super-resolution enhancement and intermediate frames that remain at lower resolution. This segmentation is based on temporal importance and motion characteristics, allowing the system to apply memory-intensive super-resolution processing only where necessary, thereby significantly reducing peak memory consumption while maintaining overall video quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by performing super-resolution enhancement on only a subset of frames (key frames) rather than all frames. This selective approach applies just enough resolution enhancement to achieve the desired visual quality for important moments, while avoiding the excessive memory consumption that would result from processing every frame at high resolution.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250238905A1Video Diffusion Model
Publication Date: 2025.07.24 GOOGLE LLC
  • US20250238905A1 patent drawing
  • US20250238905A1 patent drawing
  • US20250238905A1 patent drawing

AI summary

Provided is a video generation model for performing text-to-video (T2V) or other video generation techniques. The proposed model reduces the computational costs associated with video generation. In particular, unlike traditional T2V methods, the disclosed technology can generate the full temporal duration of a video clip at once, bypassing the need for extensive computation. As one example, a machine-learned denoising diffusion model can simultaneously process a plurality of noisy inputs that correspond to various timestamps spanning the temporal dimension of a video to simultaneously generate synthetic frames for the video that match the timestamps.