Text-Guided Video Generation Using Templates and Multimedia Segments

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video generation technologies fail to meet diverse user requirements, leading to a lack of enriched video generation manners and suboptimal user experience.

Innovation Solution

A method and apparatus for video generation that involves obtaining first text information describing video effect requirements, acquiring multimedia materials, and generating a target video that meets these requirements by combining video and image segments based on multimedia materials, using video editing templates and editing operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional video generation methods are used, then the process is simple, but the video generation manner is not enriched and user experience is suboptimal

Engineering Contradiction:
Improvevideo generation mannerVSAvoidvideo generation process
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the video generation process into multiple independent modules: text analysis module, template selection module, material processing module, and video composition module. Each module handles a specific aspect of video generation, allowing the system to support diverse video generation manners while maintaining manageable complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal video generation system that can handle multiple types of video generation requests through a single platform. The system supports various video effects, multiple multimedia material types (videos, images, audio), and different editing operations, all through a unified framework that adapts to diverse user requirements.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If video generation system supports diverse requirements, then user experience improves, but system complexity increases

Engineering Contradiction:
Improvevideo effect requirementVSAvoidgeneration system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces video editing templates as intermediary elements that mediate between user requirements and the actual video generation process. Templates pre-define complex editing operations and parameters, allowing users to achieve sophisticated video effects without directly managing the underlying system complexity. The template acts as a bridge that simplifies the interaction while supporting diverse video generation capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If manual video editing is used, then video quality can be controlled, but time consumption increases

Engineering Contradiction:
Improvevideo qualityVSAvoidvideo generation time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent pre-processes and pre-configures video editing templates with optimized parameters, editing operations, and material arrangements before actual video generation. This preliminary preparation allows the system to quickly assemble high-quality videos by applying pre-tested templates rather than requiring time-consuming manual editing for each video, thus maintaining quality while reducing generation time.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12524940B2Method, apparatus, device and storage medium for video generation
Publication Date: 2026.01.13 BEIJING ZITIAO NETWORK TECH CO LTD
  • US12524940B2 patent drawing
  • US12524940B2 patent drawing
  • US12524940B2 patent drawing

AI summary

The disclosure provides a method, an apparatus, a device and a storage medium for video generation. The method comprises: obtaining first text information used to describe a video effect requirement; obtaining at least one multimedia material; and generating a target video based on the first text information and the at least one multimedia material. The at least one multimedia material is presented in the target video. A video effect of the target video meets the video effect requirement described in the first text information. The target video is used to present a combination of at least one video segment. The at least one video segment is formed respectively based on respective video-image materials in the at least one multimedia material. The respective video-image materials comprise a video material and/or an image material.