Picture Text Translation via Paragraph Reconstruction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing need for international communication requires effective translation of words in pictures and videos to accommodate users speaking different languages, as existing methods often result in inaccurate or incomplete translations when handling multi-line text.

Innovation Solution

A method and apparatus that recognizes text lines in a picture, combines them into paragraphs, translates the paragraphs into target language, and replaces the original text with the translated paragraphs, using techniques like OCR, machine learning, and neural networks to maintain the original typesetting and improve translation accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If words in pictures are translated to accommodate users speaking different languages, then the propagation range of information is expanded, but the accuracy and completeness of multi-line text translation deteriorates

Engineering Contradiction:
Improvelanguage adaptabilityVSAvoidtranslation accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the translation process into distinct modules: text line recognition extracts individual lines from the picture, paragraph combination reconstructs logical paragraphs from segmented lines, and translation converts the reconstructed paragraphs. This segmentation allows each module to specialize, improving overall translation accuracy while maintaining language adaptability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary text line recognition and paragraph combination before translation. By pre-processing the text to reconstruct logical paragraphs from picture elements, the system prepares accurate source text for translation, ensuring both multi-line completeness and translation accuracy are maintained.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If text lines are recognized and combined into paragraphs for accurate translation, then translation completeness is improved, but the processing complexity increases

Engineering Contradiction:
Improvetranslation completenessVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent divides the complex processing task into manageable segments: text line recognition handles individual line extraction, paragraph combination handles logical structure reconstruction, and translation handles language conversion. This segmentation reduces processing complexity by making each step independent and specialized.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces paragraph combination as an intermediary step between text line recognition and translation. This mediator reconstructs logical paragraphs from recognized text lines, ensuring translation completeness while keeping the overall process structured and manageable through clear separation of concerns.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11954455B2Method for translating words in a picture, electronic device, and storage medium
Publication Date: 2024.04.09 DOUYIN VISION CO LTD
  • US11954455B2 patent drawing
  • US11954455B2 patent drawing
  • US11954455B2 patent drawing

AI summary

Provided are a method and apparatus for translating words in a picture, an electronic device, and a storage medium. The method includes: recognizing words embedded in a target picture to obtain at least one text line, each of which corresponds to one line of words; perform paragraph combination on the at least one text line to obtain at least one text paragraph; translating the at least one text paragraph into at least one target text paragraph in a specified language; and replacing the words in the target picture with the at least one target text paragraph.