Narrative-Based Video Generation System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Content producers face challenges in generating engaging and efficient video content due to the complexity, time, and cost associated with creating video content that captures user interest in a crowded market.

Innovation Solution

A content system that uses a multi-step process involving narration, image generation, animation, and text-to-speech models to automatically generate video content from input story description text, allowing for quick and efficient creation of video content by obtaining input data, generating narrative text, identifying key text subsets, producing images and video segments, and combining them with narrative speech for output.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional video content creation methods are used, then video quality and engagement can be maintained, but production time and costs increase significantly

Engineering Contradiction:
Improvevideo content generation speedVSAvoidcontent creation process complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The video content generation process is divided into distinct sequential stages: story description input → narrative text generation → image generation → video segment creation → final video assembly. Each stage is handled by a specialized model (narration model, image generation model, animation model), allowing parallel processing and optimization of individual components while maintaining overall system efficiency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate representations that bridge different processing stages: narrative text serves as an intermediary between story input and visual generation, image prompts mediate between text and visual content, and structured video segments act as intermediaries before final assembly. These intermediaries enable modular processing and facilitate the integration of multiple AI models.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of time

If traditional video production processes are followed, then content quality can be ensured, but production time and resource requirements increase

Engineering Contradiction:
Improvevideo content production timeVSAvoidcontent creation simplicity
Core Design Contradiction:
Loss of timeVSEase of manufacture

Solution Approach 1:

The system performs self-service by automatically generating narrative text from story descriptions, creating image prompts from the same input, and assembling video segments without requiring manual intervention at each stage. The AI models autonomously transform input data through multiple processing stages, eliminating the need for separate human operations in narration writing, image generation, and video editing.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The narration model and image generation model serve multiple functions: the narration model generates both the narrative text for video content and the prompts for image generation, while the image generation model creates both the visual content and associated metadata. This multi-functionality reduces the number of separate processes needed and accelerates overall production.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If automated content generation is implemented, then production efficiency increases, but control over content quality and accuracy decreases

Engineering Contradiction:
Improvecontent generation efficiencyVSAvoidcontent generation accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system incorporates feedback loops where generated narrative text is processed to create image prompts that are then used to generate visual content. The structured video segments are assembled based on feedback from the narrative analysis, ensuring consistency between the story content and visual representation. This multi-stage feedback mechanism maintains accuracy while enabling automated processing.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The narration model generates the complete narrative text and identifies key segments before visual content is created. Image prompts are prepared in advance based on the narrative analysis, and video segment structures are predetermined before actual video generation begins. These preliminary actions ensure that the subsequent automated generation processes produce accurate and coherent content.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240404163A1Video-Content System with Narrative-Based Video Content Generation Feature
Publication Date: 2024.12.05 ROKU INC
  • US20240404163A1 patent drawing
  • US20240404163A1 patent drawing
  • US20240404163A1 patent drawing

AI summary

In one aspect, an example method includes (i) obtaining input data, wherein the input data includes story description text; (ii) providing the obtained input data to a narration model and responsively receiving generated narrative text; (iii) identifying, from among the generated narrative text, a subset of text; (iv) providing the identified subset of text to an image generation model and responsively receiving generated images; (v) providing the generated images to an animation model and responsively receiving generated video segments; (vi) providing the generated narrative text to a text-to-speech model and responsively receiving generated narrative speech; (vii) combining the generated video segments and the generated narrative speech to generate video content; and (viii) outputting for presentation, the generated video content.