Synchronized Audio-Text Navigation for Document Video Playback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies do not effectively allow users to selectively listen to and reproduce specific portions of a video explanation synchronized with a document, limiting the ability to focus on particular parts of the content.

Innovation Solution

An information processing apparatus that includes a sound recognition unit, a detecting unit, an extracting unit, a display unit, and a reproducing unit, which recognizes sounds in a moving image, detects and extracts relevant words, displays them differently, and allows users to designate and reproduce the moving image from specific occurrence times, enabling focused listening and viewing of document-related content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If the entire moving image is reproduced, then the user can view all content, but the user cannot efficiently focus on specific sections of the document

Engineering Contradiction:
ImproveAbility to focus on specific contentVSAvoidTime to locate specific content
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system performs preliminary actions by extracting words from the document, recognizing sounds in the moving image, detecting matching words, and extracting their occurrence times before the user needs to search. This pre-processing creates an index that enables rapid navigation to specific content sections when users input search terms, eliminating the need to manually scan through the entire video.

Inventive Principle:
Principle #10Preliminary action

2Loss of information

If sound recognition is performed on the entire moving image, then all spoken content is captured, but the user cannot easily identify which words correspond to specific document positions

Engineering Contradiction:
ImproveCompleteness of sound recognitionVSAvoidAssociation accuracy between words and positions
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The system merges two separate data streams: sound recognition results (including occurrence times) and document word extractions (including positions). By detecting words that appear in both the recognized sound and the document, and by associating the occurrence time from sound recognition with the position from document extraction, the system creates synchronized information that maintains completeness while enabling precise location identification.

Inventive Principle:
Principle #5Merging (Combining)

3Loss of information

If the system displays all words in the document, then the user can see the complete content, but the user cannot easily identify which words are spoken in the moving image

Engineering Contradiction:
ImproveCompleteness of document displayVSAvoidIdentification of spoken words
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The system applies local quality by displaying words with different visual characteristics based on their status. Words that are detected as both in the document and in the sound recognition results are displayed with special markers (such as underlining or highlighting), while other words display normally. This selective differentiation enables users to quickly identify spoken words without losing visibility of the complete document content.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9224305B2Information processing apparatus, information processing method, and non-transitory computer readable medium storing information processing program
Publication Date: 2015.12.29 FUJIFILM BUSINESS INNOVATION CORP
  • US9224305B2 patent drawing
  • US9224305B2 patent drawing
  • US9224305B2 patent drawing

AI summary

An information processing apparatus includes a sound recognition unit that recognizes a sound of a moving image including a captured document, a detecting unit that detects a word which appears in both a recognition result of the sound recognition unit and a word extracted from the captured document in the moving image, an extracting unit that extracts an occurrence time of the word detected by the detecting unit in the moving image and a position of the word in the document, a display unit that displays the word extracted by the extracting unit on the document in a different manner from that in which another word is displayed, a designating unit that designates the word displayed by the display unit on the basis of an operation of an operator, and a reproducing unit that reproduces the moving image from the occurrence time of the word designated by the designating unit.