Synchronized Audio-Text Navigation for Document Video Playback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies do not effectively allow users to selectively listen to and reproduce specific portions of a video explanation synchronized with a document, limiting the ability to focus on particular parts of the content.
Innovation Solution
An information processing apparatus that includes a sound recognition unit, a detecting unit, an extracting unit, a display unit, and a reproducing unit, which recognizes sounds in a moving image, detects and extracts relevant words, displays them differently, and allows users to designate and reproduce the moving image from specific occurrence times, enabling focused listening and viewing of document-related content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If the entire moving image is reproduced, then the user can view all content, but the user cannot efficiently focus on specific sections of the document
Solution Approach 1:
The system performs preliminary actions by extracting words from the document, recognizing sounds in the moving image, detecting matching words, and extracting their occurrence times before the user needs to search. This pre-processing creates an index that enables rapid navigation to specific content sections when users input search terms, eliminating the need to manually scan through the entire video.
2Loss of information
If sound recognition is performed on the entire moving image, then all spoken content is captured, but the user cannot easily identify which words correspond to specific document positions
Solution Approach 1:
The system merges two separate data streams: sound recognition results (including occurrence times) and document word extractions (including positions). By detecting words that appear in both the recognized sound and the document, and by associating the occurrence time from sound recognition with the position from document extraction, the system creates synchronized information that maintains completeness while enabling precise location identification.
3Loss of information
If the system displays all words in the document, then the user can see the complete content, but the user cannot easily identify which words are spoken in the moving image
Solution Approach 1:
The system applies local quality by displaying words with different visual characteristics based on their status. Words that are detected as both in the document and in the sound recognition results are displayed with special markers (such as underlining or highlighting), while other words display normally. This selective differentiation enables users to quickly identify spoken words without losing visibility of the complete document content.
Data Source
AI summary
An information processing apparatus includes a sound recognition unit that recognizes a sound of a moving image including a captured document, a detecting unit that detects a word which appears in both a recognition result of the sound recognition unit and a word extracted from the captured document in the moving image, an extracting unit that extracts an occurrence time of the word detected by the detecting unit in the moving image and a position of the word in the document, a display unit that displays the word extracted by the extracting unit on the document in a different manner from that in which another word is displayed, a designating unit that designates the word displayed by the display unit on the basis of an operation of an operator, and a reproducing unit that reproduces the moving image from the occurrence time of the word designated by the designating unit.


