Template-Guided Video Synthesis Using Generative AI

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video synthesis techniques using deep neural networks generate videos of lower quality due to lack of temporal consistency, while template-based approaches are constrained by limited and fixed image and text inputs.

Innovation Solution

A method that combines generative artificial intelligence models with template-guided video synthesis, where text content is generated based on user or sponsor information and applied to a template model along with image content to produce high-quality video.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If deep neural networks are used to generate video content, then generation speed and efficiency are improved, but video quality and temporal consistency deteriorate

Engineering Contradiction:
Improvevideo generation speedVSAvoidvideo quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The video generation process is segmented into multiple stages: first generating text content from input information using a generative AI model, then using that text along with images as inputs to a template model to generate the final video. This segmentation allows each stage to optimize for its specific function, with the template model ensuring temporal consistency in the video generation phase.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Text content serves as an intermediary element between the input information and the final video output. The generative AI model first produces text content from input information, which then serves as a prompt or guide for the template model to generate the video, thereby mediating between rapid generation needs and quality requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If template-based approaches are used to generate video content, then temporal consistency is improved, but content diversity and adaptability deteriorate

Engineering Contradiction:
Improvetemporal consistencyVSAvoidcontent diversity
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The system dynamically adapts the template model inputs based on generative AI-generated text content. Instead of using fixed inputs, the text content generated from input information serves as a dynamic prompt that guides the template model to create diverse and customized video content while maintaining temporal consistency through the template structure.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The template model serves multiple functions: it provides temporal consistency through its structured approach, enables content customization through dynamic text prompts, and maintains adaptability by processing varied input information. This multi-functionality resolves the contradiction between consistency and diversity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If fixed template models are used, then production control is improved, but creativity and customization deteriorate

Engineering Contradiction:
Improveproduction controlVSAvoidcustomization capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary action by generating text content from input information before feeding it to the template model. This preliminary text generation step enables the template model to work with customized, dynamically generated prompts rather than fixed inputs, thereby maintaining production control while enhancing customization capability.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250133273A1Machine learning assisted and template guided video synthesis
Publication Date: 2025.04.24 GOOGLE LLC
  • US20250133273A1 patent drawing
  • US20250133273A1 patent drawing
  • US20250133273A1 patent drawing

AI summary

A method for generating video content includes obtaining first information that includes information associated with a user or a content sponsor. The method also includes generating text content at least in part by applying the first information to a generative artificial intelligence model, and obtaining image content. The method further includes generating video content, at least by applying the text content and the image content as inputs to a template model. The template model causes the generated video content to conform to one or more temporal characteristics defined by the template model.