Narrative-Based Video Content Generation System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge for content producers is to generate engaging video content efficiently and cost-effectively, as the abundance of video content makes it difficult to capture users' attention, and traditional methods are time-consuming and expensive.

Innovation Solution

A content system that takes story description text as input, uses a narration model to generate narrative text, identifies a subset for image generation, produces images and video segments through an animation model, and combines this with text-to-speech-generated audio to create synthetic video content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional video production methods are used, then video content quality can be maintained, but production time and costs increase significantly

Engineering Contradiction:
Improvevideo content generation speedVSAvoidproduction time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces traditional mechanical video production processes (filming, editing, rendering) with AI-based generative models. The system uses text-to-video generation models, image generation models, and animation models to automatically create video content from text descriptions, eliminating the need for manual production workflows and significantly reducing production time

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service video content generation where users can input text descriptions and the AI models automatically generate complete video content without requiring professional video production expertise. The workflow includes automatic narrative generation, image creation, animation, and audio synthesis, allowing users to independently produce video content

Inventive Principle:
Principle #25Self-service

2Productivity

If traditional video production methods are used, then video content can be created, but production costs increase significantly

Engineering Contradiction:
Improvevideo content generation efficiencyVSAvoidproduction cost
Core Design Contradiction:
ProductivityVSEase of manufacture

Solution Approach 1:

The patent uses generative AI models to create synthetic video content that copies and transforms text descriptions into visual and audio representations. The system generates narrative text from input descriptions, creates images from text, and synthesizes audio from text, replacing expensive production resources with computational generation

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system employs multi-functional AI models that can perform multiple production tasks. The same generative infrastructure handles text-to-video, text-to-audio, image generation, and animation, consolidating multiple specialized production tools into a unified system that reduces overall production costs

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If abundant video content is produced, then user choice increases, but capturing user attention becomes more difficult

Engineering Contradiction:
Improvevideo content engagementVSAvoiduser attention
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system enhances video content engagement by generating high-quality, localized visual and audio elements that specifically match the narrative content. The image generation model creates visually striking images, the animation model adds dynamic motion, and the text-to-speech model provides expressive narration, ensuring each element contributes to capturing user attention

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12039653B1Video-content system with narrative-based video content generation feature
Publication Date: 2024.07.16 ROKU INC
  • US12039653B1 patent drawing
  • US12039653B1 patent drawing
  • US12039653B1 patent drawing

AI summary

In one aspect, an example method includes (i) obtaining input data, wherein the input data includes story description text; (ii) providing the obtained input data to a narration model and responsively receiving generated narrative text; (iii) identifying, from among the generated narrative text, a subset of text; (iv) providing the identified subset of text to an image generation model and responsively receiving generated images; (v) providing the generated images to an animation model and responsively receiving generated video segments; (vi) providing the generated narrative text to a text-to-speech model and responsively receiving generated narrative speech; (vii) combining the generated video segments and the generated narrative speech to generate video content; and (viii) outputting for presentation, the generated video content.