Text-to-Image Video Creation With Interactive Storytelling Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text-to-image generation applications lack user interaction and are time-consuming for users trying to create suitable images for content creation, particularly for generating storytelling videos, leading to a lack of creativity and user engagement.

Innovation Solution

A system utilizing machine learning models to generate images from text prompts, allowing users to create interactive storytelling videos by blending generated images with live camera feeds, text graphics, and pre-recorded audio, enabling multiple image generation from a series of sentences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If users manually search for or create images from scratch, then they can find suitable images for content creation, but it is time-consuming and reduces productivity

Engineering Contradiction:
Improveimage suitabilityVSAvoidcontent creation speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent replaces manual image searching and creation processes with an automated text-to-image generation system using machine learning models. Users input text prompts describing desired images, and the system automatically generates suitable images, eliminating the time-consuming manual processes while maintaining image quality and relevance.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If existing text-to-image applications generate images automatically, then productivity is improved, but user interaction and creativity are reduced

Engineering Contradiction:
Improveimage generation speedVSAvoiduser interaction
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The system incorporates iterative feedback mechanisms where users can review generated images and provide corrections or refinements to their text prompts. This allows users to maintain creative control and interact with the generation process, ensuring the final images match their vision while still benefiting from automated generation speed.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent implements dynamic user interaction where users can adjust parameters, refine prompts, and iteratively improve generated images. The system adapts to user preferences and provides real-time feedback, transforming a static automated process into a dynamic collaborative creation experience.

Inventive Principle:
Principle #15Dynamics

3Manufacturing precision

If high-resolution images are generated, then image quality is improved, but waiting time increases

Engineering Contradiction:
Improveimage resolutionVSAvoidgeneration waiting time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent divides the image generation process into multiple stages or segments, allowing users to view progressive results. The system can generate lower-resolution previews quickly for immediate feedback, then progressively refine to higher resolutions, reducing perceived waiting time while maintaining final image quality.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12482497B2Content creation based on text-to-image generation
Publication Date: 2025.11.25 LEMON INC(GB)
  • US12482497B2 patent drawing
  • US12482497B2 patent drawing
  • US12482497B2 patent drawing

AI summary

The present disclosure describes techniques for generating content. Text may be received. The text is associated with a video to be created by at least one user. At least one image may be generated based at least in part on the text using at least one machine learning model. The video may be generated based at least in part on the at least one image. The video comprises content overlaid on the at least one image.