Video Manual Generation Linking Text Nouns to Video Actions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video manual generation methods struggle to clearly convey the correspondence between objects and actions, making it difficult for users to understand the intended procedures.
Innovation Solution
A device and method that analyzes work procedure manuals and video recordings to generate text and video data, linking object positions and actions by collecting and matching noun-verb sets from text and video data, and displaying this information synchronously in a video manual to enhance understanding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If video manual is generated by analyzing video and generating event list with time information and object information, then the video manual can be created as a multimedia manual, but the correspondence between object and action of person is hard to understand
Solution Approach 1:
The patent segments the video analysis into distinct components: object detection (identifying tools and components), action detection (identifying worker movements and operations), and event generation (combining object and action with time information). This segmentation allows each component to be optimized independently while maintaining their relationships through the event structure, thereby preserving object-action correspondence clarity while enabling efficient multimedia manual generation.
2Device complexity
If text information and video information are processed separately, then the processing complexity is reduced, but the correspondence between text description and video content is difficult to establish
Solution Approach 1:
The patent introduces an event list as an intermediary data structure that bridges text information and video information. The event list contains time information, object information, and action information that are extracted from video analysis, and this structured event data serves as the foundation for generating both the video manual and the text description. This intermediary ensures consistent and accurate correspondence between text and video content while keeping the processing system manageable through standardized data formats.
Data Source
AI summary
A video manual generation device includes a document analysis unit, a video analysis unit, a link information generation unit to collect first sets each being a set of a noun and a verb from text information data, to collect second sets each being a set of an object and an action from object information data and action information data, to search those sets for a first set and a second set in which the noun and the object correspond to each other and the verb and the action correspond to each other, and to generate link information data indicating correspondence between a position in the work procedure where the first set obtained by the search is described and a scene in a video that includes the second set obtained by the search, and a video manual generation unit to generate video manual data based on the link information data.


