Subtitle Generation Apparatus for Automated Video Text Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies for generating translated subtitles from videos in foreign languages are inefficient, requiring significant time and effort, as they rely on manual processes and lack automation for text extraction and translation.

Innovation Solution

A subtitle generation apparatus comprising a text information extraction unit, text coincidence detection unit, text translation unit, display position calculation unit, and subtitle synthesizing unit, which automatically extracts text from video data, detects relevant dialogue information, translates text, calculates optimal display positions, and synthesizes translated subtitles into the video.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual processes are used to generate translated subtitles from video data, then translation accuracy can be maintained, but significant time and effort are required

Engineering Contradiction:
Improvetranslation accuracyVSAvoidsubtitle generation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces an automatic text extraction unit and translation unit as intermediaries between the video data and the final subtitles. These automated components process the video data to extract text and translate it, significantly reducing the time required while maintaining acceptable accuracy through subsequent manual review and editing processes.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The subtitle generation process is divided into distinct segments: automatic text extraction from video, automatic translation of extracted text, manual review and editing, and final subtitle generation. This segmentation allows different levels of automation to be applied to different stages, optimizing both speed and accuracy.

Inventive Principle:
Principle #1Segmentation

2Quantity of substance

If all text in video data is translated, then completeness of translation is improved, but unnecessary translation of non-dialogue text increases work effort

Engineering Contradiction:
Improvecompleteness of translationVSAvoidwork effort
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent applies different processing qualities to different types of text in the video. Dialogue text is automatically extracted and translated with high priority, while other text elements are either excluded or processed differently. This local differentiation of translation quality based on text type and importance reduces unnecessary work effort while maintaining completeness for essential content.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

Instead of translating all text equally, the system performs partial action by selectively translating only the dialogue portions that are most important for subtitle generation. This approach translates sufficient content to maintain completeness for the intended purpose while avoiding excessive translation of non-essential text elements.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11363217B2Subtitle generation apparatus, subtitle generation method, and non-transitory storage medium
Publication Date: 2022.06.14 JVC KENWOOD CORP
  • US11363217B2 patent drawing
  • US11363217B2 patent drawing
  • US11363217B2 patent drawing

AI summary

An apparatus includes a text information extraction unit that extracts character information from a video including characters, a text coincidence detection unit that detects character information included in dialogue information that is data of a dialogue associated with the video data from the extracted character information, a text translation unit that translates the character information, a display position calculation unit that calculates a display position of translated text information in the video data on the basis of a text region information that indicates a region in which a video corresponding to the character information is displayed in the video data and on the basis of the translated text information, and a subtitle synthesizing unit that adds, as a subtitle, the translated text information on the basis of display position information.