Natural-Language Image Training Data from Video–Script Pairing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional training data for image generation models, particularly from video and script files, suffer from low alignment due to insufficient search information and text data, and challenges in accurately matching script elements to video content, such as scene transitions and dialogues.

Innovation Solution

A method and device that processes video and script files to generate natural language-based training data by selecting representative images, applying filters, extracting and encoding text, and pairing images with script data to create a dataset, enhancing alignment and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If web scraping is used to collect search information and text data, then the quantity of training data is increased, but the alignment between text data and images deteriorates

Engineering Contradiction:
Improvequantity of training dataVSAvoidalignment between text and image
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent extracts only the necessary text information from video transcripts and pairs it with corresponding video frames. By taking out only the relevant textual descriptions and matching them with their corresponding visual frames, the system maintains high alignment while still building a substantial training dataset from video content.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses video frames as an intermediary between the text transcript and the final training data. The video frame serves as a mediator that temporally aligns the text description with the corresponding visual content, ensuring accurate pairing while enabling large-scale data collection from video sources.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If script data is used to describe video content, then the detail of text information is improved, but the accuracy of matching script to video deteriorates

Engineering Contradiction:
Improvedetail of text informationVSAvoidaccuracy of script to video matching
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent segments the script data into scene-level descriptions and pairs them with corresponding video frames. By dividing the continuous script into discrete segments that match specific video moments, the system maintains detailed textual information while improving the precision of script-video alignment through temporal segmentation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary processing of the script data to identify and extract only the relevant descriptive portions that correspond to actual video content. This preliminary action filters out unrelated script elements before pairing, ensuring high accuracy in script-video matching while preserving detailed information about the video content.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250272789A1Method and device for generating natural language -based training data using video and script files, and performing image generation and inference using the training data
Publication Date: 2025.08.28 CJ OLIVENETWORKS
  • US20250272789A1 patent drawing
  • US20250272789A1 patent drawing
  • US20250272789A1 patent drawing

AI summary

The present invention relates to a method and device for generating natural language-based training data by pairing video and script files, and performing image search and inference using the generated data including images based on input text prompts.