Comic Panel Segmentation and ML Prompting for Motion Rendering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current digital formats of graphic narratives, such as comic books, fail to leverage advancements in artificial intelligence and machine learning to provide an immersive and interactive user experience.

Innovation Solution

A method and system that utilizes machine learning models to convert static comic book panels into moving pictures by segmenting elements, applying labels, and generating prompts for a second ML model to output a moving picture, incorporating script and storyboard information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If machine learning models are used to convert static comic book panels into moving pictures, then user experience becomes more immersive and interactive, but device complexity increases

Engineering Contradiction:
Improveuser experience immersionVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments the comic book panel conversion process into distinct functional modules: a first ML model for labeling segmented elements (images and text), a prompt generation component, and a second ML model for generating moving pictures. This segmentation allows each component to be optimized independently while working together to achieve immersive user experience.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary prompt generation system that translates labels from the first ML model into instructions for the second ML model. This intermediary layer enables seamless integration between different ML models and formats, managing system complexity while maintaining versatility.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If multiple ML models are applied to generate moving pictures from panels, then conversion quality improves, but processing time increases

Engineering Contradiction:
Improveconversion qualityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary segmentation and labeling of comic book panels using the first ML model before generating moving pictures. By pre-processing the panels into labeled elements (images and text), the system prepares data in advance, enabling the second ML model to generate high-quality moving pictures more efficiently without redundant processing steps.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent establishes a continuous workflow where the output of the first ML model feeds directly into prompt generation, which then feeds into the second ML model for moving picture generation. This continuous action minimizes idle time and maintains processing momentum, reducing overall processing time while preserving conversion quality.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS20250265758A1Automated conversion of comic book panels to motion-rendered graphics
Publication Date: 2025.08.21 GLOBAL PUBLISHING INTERACTIVE INC
  • US20250265758A1 patent drawing
  • US20250265758A1 patent drawing
  • US20250265758A1 patent drawing

AI summary

A system and method are provided for generating a moving picture from a graphic narrative. Pages of a graphic narrative (e.g., comic book) are partitioned into panels, which are segmented into image segmented elements and text elements. The segmented elements are applied to a machine learning (ML) method that labels/identifies the segmented elements. Prompts based on the labels are then applied to a second ML model that outputs a moving picture representing one or more of the panels. The prompts can include script information, such as a script, storyboard, or a scene (e.g., keyframes). Thus, the comic book is effectively a movie storyboard that is automatically converted into full-motion rendered graphics by treating each combination of text and graphics as a unique prompt for a generative ML model.