Subtitle Generating Apparatus Using Archive-Based Text Splitting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional subtitle generating apparatuses face a significant burden in correcting subtitles for improved readability, especially when using real-time voice recognition results, due to errors caused by background sounds, technical terminology, and other factors.
Innovation Solution
A subtitle generating apparatus that includes processing circuitry to acquire and store voice recognition results as archive datasets, estimate split and concatenation positions, and generate subtitle texts based on these positions, thereby reducing the need for manual corrections.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If accelerated confirmation is used to terminate voice recognition search early, then real-time subtitle generation is enabled, but voice recognition results are cut short requiring manual correction
Solution Approach 1:
The system performs preliminary actions by storing multiple voice recognition results (including accelerated and non-accelerated versions) in an archive before subtitle generation. This allows the subtitle generating apparatus to select and combine results optimally without waiting for complete recognition, thus maintaining speed while improving accuracy through pre-prepared alternatives.
Solution Approach 2:
The voice recognition process is segmented into multiple results with different completion levels. The subtitle generating apparatus divides the selection process into acquiring individual results, storing them separately in the archive, and then selectively combining them. This segmentation allows optimal balance between speed and accuracy for each segment.
2Manufacturing precision
If manual correction is performed to improve subtitle readability, then subtitle quality increases, but labor burden and time consumption increase
Solution Approach 1:
The system implements self-service by automatically generating readable subtitles without requiring manual correction. The subtitle generating apparatus autonomously acquires voice recognition results from the archive, estimates appropriate split and concatenation positions, and generates final subtitles with proper formatting. This eliminates the need for human operators to perform time-consuming manual corrections while maintaining high readability.
Solution Approach 2:
The manual correction process is replaced by an automated mechanical system. The subtitle generating apparatus uses processing circuitry to automatically perform tasks that previously required human operators: selecting results from the archive, estimating split positions, concatenating texts, and formatting subtitles. This substitution eliminates labor burden and correction time while maintaining or improving subtitle quality.
3Productivity
If voice recognition search space is narrowed to improve speed, then real-time processing is achieved, but recognition accuracy decreases requiring more manual work
Solution Approach 1:
The system merges multiple voice recognition results from the archive, combining the strengths of different recognition approaches. By storing both accelerated confirmation results (for speed) and non-accelerated results (for accuracy) in the same archive and selectively combining them, the system achieves both high productivity and high reliability without requiring manual work to correct errors.
Data Source
AI summary
According to one embodiment, a subtitle generating apparatus includes processing circuitry and a display. The processing circuitry is configured to sequentially acquire texts from voice recognition results. The processing circuitry is configured to store the texts as archive datasets. The processing circuitry is configured to estimate a split position and a concatenation position of the texts from one or more of the archive datasets, and generate a subtitle text from said one or more of the archive datasets based on the split position and the concatenation position. The processing circuitry is configured to update the archive datasets based on the split position and the concatenation position. The display is configured to display the subtitle text.


