Video Manual Generation Using Task-Trained Model for Procedure Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video manual generation apparatuses face high processing loads due to the need to recognize combinations of objects and actions in videos, and they struggle to uniquely associate scenes with procedures when these combinations occur multiple times.

Innovation Solution

A video manual generation apparatus that utilizes a task-trained model to identify procedures corresponding to frames in input video data, reducing processing load and enabling accurate association of scenes with procedures even when object-action combinations recur.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional object-action recognition is used to associate video scenes with procedures, then the association can be made, but the processing load becomes heavy

Engineering Contradiction:
Improveassociation accuracyVSAvoidprocessing load
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent changes the recognition parameters from detailed object-action combinations to simpler procedure-level identification. The system uses procedure IDs and temporal information instead of complex object and action recognition, significantly reducing processing load while maintaining association accuracy.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent extracts only the essential information needed for association - procedure identification and temporal positioning - while discarding the complex object-action recognition process. This extraction approach maintains the core functionality while reducing processing complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If object-action recognition is used, then video scenes can be identified, but unique association becomes impossible when combinations recur

Engineering Contradiction:
Improveunique association capabilityVSAvoidprocedure identification accuracy
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent adds temporal dimension to the association process. By incorporating time information and procedure sequence, the system can uniquely identify procedures even when object-action combinations repeat. The association is no longer based solely on content similarity but also on temporal positioning and procedure context.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces procedure ID as an intermediary element between video scenes and procedures. This intermediary allows unique identification and association even when the actual object-action content repeats, solving the ambiguity problem.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250157216A1Video manual generation apparatus
Publication Date: 2025.05.15 NTT DOCOMO INC
  • US20250157216A1 patent drawing
  • US20250157216A1 patent drawing
  • US20250157216A1 patent drawing

AI summary

A video manual generation apparatus includes an acquirer configured to acquire: an input-video data set indicative of contents of a task including one or more procedures, and one or more input-procedure-text data sets in one-to-one correspondence with the one or more procedures, an identifier configured to use a task trained model to identify a procedure corresponding to a frame that is any one of a plurality of frames of the input-video data set from among the one or more procedures, the task trained model being trained to learn a relationship between first information and second information, the first information being constituted of a video and one or more texts, the video representing the contents of the task constituted of the one or more procedures, the one or more texts being in one-to-one correspondence with the one or more procedures, the second information indicating a procedure corresponding to a frame that is any one of a plurality of frames of the video among the one or more procedures; and a video manual generator configured to generate video manual data based on the input-video data set and an input-procedure-text data set corresponding to the procedure identified by the identifier from among the one or more input-procedure-text data sets.