Synchronized Video and Text Playback for Conversation Review
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Checking specific utterances in video files with voice recordings is cumbersome and labor-intensive, as it requires synchronizing visual and audio elements, placing a heavy burden on reviewers.
Innovation Solution
A video playback system that includes a video acquisition unit, a text acquisition unit, and display units to record and playback video and text data synchronously, allowing for the selection and display of text data corresponding to voice recognition at specific image capture times, facilitating easier review of face-to-face interactions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If video files with voice recordings are used to record face-to-face interactions, then the interaction can be captured comprehensively, but checking specific utterances becomes labor-intensive and time-consuming
Solution Approach 1:
The patent segments the continuous video and voice data by synchronizing it with transcribed text data. This allows the checker to navigate to specific utterances by locating corresponding text segments rather than reviewing entire video files, significantly reducing checking time while maintaining recording completeness
Solution Approach 2:
The patent introduces text data (transcriptions) as an intermediary between the video/voice recording and the checking process. This text intermediary enables rapid location and verification of specific utterances without requiring direct analysis of audio or video content, resolving the contradiction between comprehensive recording and efficient checking
2Reliability
If video files with voice recordings are used to record face-to-face interactions, then the interaction can be captured comprehensively, but the checking process places a heavy burden on reviewers
Solution Approach 1:
By dividing the continuous media into discrete text-synchronized segments, the system allows reviewers to focus on specific utterances of interest rather than processing entire video files, reducing cognitive and operational burden
Solution Approach 2:
The patent replaces the mechanical process of listening to and watching video/voice content with a more efficient text-based search and review process. Reviewers can quickly scan, search, and locate specific utterances in text form, significantly reducing the operational burden compared to traditional video review methods
3Productivity
If text data is synchronized with video data at image capture timing, then specific utterances can be located efficiently, but the system complexity increases
Solution Approach 1:
The patent creates a text copy (transcription) of the spoken content and synchronizes it with the video data using timestamps. This text copy serves as a searchable and navigable representation that enables efficient location of utterances without requiring complex analysis of the original video and audio streams
Solution Approach 2:
The patent uses timing parameters (timestamps indicating image capture timing) to synchronize text data with video data. By matching the timing parameter of when each frame was captured with corresponding text segments, the system enables efficient utterance location through simple time-based correlation rather than complex multi-modal analysis
Data Source
AI summary
An after-the-fact check of a face-to-face interaction between a plurality of people with a conversation. A video playback system includes a video acquisition unit, a video memory unit, a text acquisition unit, a text memory unit, a video display unit, and a text display unit. The video memory unit stores video data representing a video captured over an image capture area and acquired by the video acquisition unit. The text memory unit stores text data generated by recognizing a voice collected around the image capture area and acquired by the text acquisition unit. The text display unit selects text data formed by recognizing a voice collected around the image capture area at an image capture timing of the video displayed by the video display unit, from among the stored text data, and displays a text represented by the text data.


