Voice-Based Video Indexing Using Language Models for Skill Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems fail to accurately associate moving image data with index information that represents the content of the data, making it difficult for unskilled workers to efficiently search and learn from captured work performances by skilled workers.
Innovation Solution
An information processing system that converts voice data into character strings, extracts relevant words using a language learning model like BERT, and associates these words with moving image data to generate accurate index information, enabling effective search and retrieval of relevant content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing systems store moving image data with basic index information, then storage is simple, but the index information does not accurately represent the content of the data
Solution Approach 1:
The patent introduces voice data as an intermediary element that bridges moving image data and index information. The system captures voice data alongside moving image data, uses it to generate accurate index information through text conversion and word extraction, and stores all three elements in association. This intermediary approach enables content-accurate indexing without requiring direct complex analysis of the moving image data itself.
Solution Approach 2:
The patent replaces manual or simple automatic indexing methods with an automated language processing system. By using voice data conversion and automated word extraction techniques, the system substitutes complex manual indexing work with automated computational processes, achieving both high accuracy and operational efficiency.
2Productivity
If manual indexing is used for moving image data, then system complexity is low, but search efficiency and accuracy deteriorate
Solution Approach 1:
The system enables self-service indexing by automatically generating index information from voice data that is naturally captured during work activities. The voice data serves as self-generated metadata that automatically describes the moving image content, eliminating the need for external manual indexing while maintaining high search efficiency and accuracy.
Solution Approach 2:
The system performs preliminary indexing by capturing and processing voice data at the time of work execution. This preliminary action creates accurate index information in advance, enabling efficient subsequent searches without requiring complex real-time analysis during retrieval operations.
3Measurement precision
If voice data is converted and processed to generate index information, then indexing accuracy improves, but processing time increases
Solution Approach 1:
The system performs voice-to-text conversion and word extraction as preliminary actions during or immediately after work execution. By completing the indexing process in advance rather than during search operations, the system achieves high indexing accuracy while minimizing the time loss impact on overall workflow efficiency.
Data Source
AI summary
An information processing method, by a processing unit of an information processing apparatus, includes: converting voice data into character string data; generating question data by extracting a first word from the character string data; extracting a second word from the character string data by inputting the character string data and the question data to a trained language learning model configured to output, when the character string data and the question data are input, a word corresponding to an answer to the question data from the character string data; and storing the voice data, the first word, and the second word in association with each other.


