Visual Work Instruction Search Using Segmented Video Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing AI-based search models face challenges in handling large document lengths and require extensive labeled data for training, while unsupervised learning models struggle with long documents and non-textual data like videos and images.
Innovation Solution
A method and apparatus for generating segment search data by separating video segments based on textual information, creating a text file, and storing synchronization information to enable AI-based text search in visual work instructions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If AI-based search models are used to improve search accuracy, then search precision is improved, but the cost of data labeling increases significantly
Solution Approach 1:
The patent segments long videos into smaller clips based on textual information and timestamps, creating manageable units for training AI models. This segmentation reduces the overall complexity and cost of labeling while maintaining search accuracy through targeted training on representative segments.
Solution Approach 2:
The patent performs preliminary extraction of textual information from videos and creates segment metadata before AI model training. This preliminary action prepares structured data that reduces the need for extensive manual labeling during the training phase, lowering costs while preserving search accuracy.
2Loss of information
If AI-based search models process long documents, then comprehensive search coverage is improved, but the model's processing capability deteriorates due to token limits
Solution Approach 1:
The patent divides long videos into multiple segments with associated textual metadata, allowing the AI model to process each segment within token limits while maintaining comprehensive search coverage across the entire video through aggregated segment indexing.
Solution Approach 2:
The patent transitions from processing entire long videos in a single dimension to processing segmented clips with textual metadata in multiple dimensions (video segments, text descriptions, timestamps), enabling comprehensive coverage while respecting model token limitations.
3Adaptability or versatility
If unsupervised learning models are used to handle non-textual data, then data processing versatility is improved, but search precision deteriorates
Solution Approach 1:
The patent introduces textual information and metadata as an intermediary between non-textual video data and the AI search model. This intermediary layer enables unsupervised learning models to process diverse video content while maintaining search precision through structured text-based queries and segment matching.
Data Source
AI summary
The present invention relates to generating training data for performing artificial intelligence. A method and apparatus for generating section data of a visual work instruction for performing artificial intelligence are provided for generating data for searching a user's desired section in a visual work instruction using an artificial intelligence-based text search model.


