Display Timing Determination for Voice-Subtitle Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies fail to accurately synchronize the output timing of voices with the display timing of character information, such as subtitles, especially when genre codes are not available, leading to timing delays and inaccuracies due to variations in content complexity and editor skill.
Innovation Solution
A system that includes voice storage data acquisition, timing data acquisition, waveform analysis, and display timing determination to synchronize voice and character information display by analyzing voice waveforms and adjusting provisional display timings based on match degree information to determine definitive display timings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual character creation is used in live TV programs, then character information can be provided, but display timing is delayed relative to voice output timing
Solution Approach 1:
The system uses automatic waveform analysis to extract timing information directly from the audio signal itself, eliminating the need for manual timing adjustment. The waveform analysis means automatically determines when characters should be displayed based on the voice waveform patterns, making the system self-sufficient without requiring manual intervention for timing synchronization.
Solution Approach 2:
The patent replaces the manual mechanical process of timing adjustment with automated computational waveform analysis. Instead of manually calculating and adjusting timing based on genre codes or editor experience, the system uses digital signal processing to automatically extract timing information from the audio waveform, substituting human expertise with automated algorithms.
2Adaptability or versatility
If delay time is estimated based on genre code, then display timing can be adjusted, but synchronization accuracy deteriorates due to variations in content complexity and editor skill
Solution Approach 1:
The system determines timing adjustments individually for each character based on its specific waveform characteristics rather than applying a uniform delay time for all characters in a genre. The waveform analysis means examines the local waveform patterns of each voice segment to determine precise timing, allowing each character to have optimized display timing based on its actual acoustic properties.
Solution Approach 2:
The system uses the actual waveform data as feedback to continuously refine timing determination. By analyzing the real waveform patterns and comparing them against the character display timing, the system can adjust and optimize synchronization accuracy based on the actual content characteristics, creating a feedback-driven optimization process.
3Productivity
If provisional display timings are used, then character information can be displayed, but timing precision is insufficient without waveform analysis
Solution Approach 1:
The system first establishes provisional display timings based on the character information and voice output sequence, then uses waveform analysis to refine and adjust these provisional timings to achieve precise synchronization. This two-stage approach allows for efficient initial timing setup followed by precise optimization based on actual waveform characteristics.
Data Source
AI summary
Voice storage data acquisition means of a display timing determination device acquires voice storage data storing a plurality of voices to be output sequentially. Timing data acquisition means acquires timing data on provisional display timings of a plurality of pieces of character information, which are to be sequentially displayed during reproduction of the voice storage data and represent content of the respective voices. Waveform analysis means analyzes a voice waveform of the voice storage data to acquire an output timing of each of the voices. Display timing determination means determines a definitive display timing of each of the pieces of character information based on the output timing of each of the voices acquired by the waveform analysis means and the provisional display timings of the respective pieces of character information determined based on the timing data.


