Natural-Language Image Training Data from Video–Script Pairing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional training data for image generation models, particularly from video and script files, suffer from low alignment due to insufficient search information and text data, and challenges in accurately matching script elements to video content, such as scene transitions and dialogues.
Innovation Solution
A method and device that processes video and script files to generate natural language-based training data by selecting representative images, applying filters, extracting and encoding text, and pairing images with script data to create a dataset, enhancing alignment and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If web scraping is used to collect search information and text data, then the quantity of training data is increased, but the alignment between text data and images deteriorates
Solution Approach 1:
The patent extracts only the necessary text information from video transcripts and pairs it with corresponding video frames. By taking out only the relevant textual descriptions and matching them with their corresponding visual frames, the system maintains high alignment while still building a substantial training dataset from video content.
Solution Approach 2:
The patent uses video frames as an intermediary between the text transcript and the final training data. The video frame serves as a mediator that temporally aligns the text description with the corresponding visual content, ensuring accurate pairing while enabling large-scale data collection from video sources.
2Loss of information
If script data is used to describe video content, then the detail of text information is improved, but the accuracy of matching script to video deteriorates
Solution Approach 1:
The patent segments the script data into scene-level descriptions and pairs them with corresponding video frames. By dividing the continuous script into discrete segments that match specific video moments, the system maintains detailed textual information while improving the precision of script-video alignment through temporal segmentation.
Solution Approach 2:
The patent performs preliminary processing of the script data to identify and extract only the relevant descriptive portions that correspond to actual video content. This preliminary action filters out unrelated script elements before pairing, ensuring high accuracy in script-video matching while preserving detailed information about the video content.
Data Source
AI summary
The present invention relates to a method and device for generating natural language-based training data by pairing video and script files, and performing image search and inference using the generated data including images based on input text prompts.


