Image Text Translation via OCR and Bounding Box Overlay

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face inconvenience when translating text from images, especially in foreign languages, as they need to manually input text, which is cumbersome and inconvenient, especially when traveling or dealing with unfamiliar scripts like Arabic or Chinese.

Innovation Solution

A method for translating text in images by capturing and processing images to extract and translate text from a source language to a destination language, allowing users to view the translated text within the same image bounding boxes, with options for font size adjustment and user-defined gestures for switching between languages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If users manually input text for translation, then translation accuracy is improved, but ease of operation deteriorates

Engineering Contradiction:
Improvetranslation accuracyVSAvoidease of operation
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs automatic text extraction and translation without requiring user input. The device captures images, automatically extracts text using OCR, translates it, and presents the results, allowing the system to serve itself rather than requiring manual user intervention for text input

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual text input with automated optical character recognition (OCR) technology. The system uses image processing and machine learning algorithms to automatically extract and translate text from captured images, substituting the manual typing mechanism with an automated visual recognition system

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If text is extracted and translated from images, then ease of operation is improved, but device complexity increases

Engineering Contradiction:
Improveease of operationVSAvoiddevice complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The translation device is designed to perform multiple functions: capturing images, extracting text through OCR, translating text between languages, and displaying results. This multi-functional approach consolidates what could be separate devices into one universal translation system, managing complexity through integration rather than proliferation of components

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system embeds multiple processing layers within each other: image capture contains text extraction, which contains translation processing, which contains display generation. Each function is nested within the previous one, creating a compact hierarchical structure that manages complexity through organized nesting rather than separate independent systems

Inventive Principle:
Principle #7Nested doll (Nesting)

3Manufacturing precision

If font sizes are adjusted for different words, then manufacturing precision is improved, but device complexity increases

Engineering Contradiction:
Improvelayout precisionVSAvoiddevice complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system applies different font sizes to different words or text regions based on their specific characteristics, such as importance, length, or original positioning in the source image. This local differentiation improves the visual quality and readability of the translated output without requiring a complete redesign of the entire display system

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11593570B2System and method for translating text
Publication Date: 2023.02.28 CONSUMER LEDGER INC
  • US11593570B2 patent drawing
  • US11593570B2 patent drawing
  • US11593570B2 patent drawing

AI summary

The subject matter discloses a method for translating text in an image, comprising extracting at least a portion of the text in a source language from the image, identifying one or more bounding boxes containing the text in the image, translating at least a portion of the text in the source language to a destination language, generating a new image containing the text in the destination language in the bounding boxes of the associated words in the source language.