Media Text Translation Using Selected-Region OCR
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing language translation methods, such as dictionaries and electronic translators, are limited to specific languages and require multiple devices for travel across different countries, lacking universality.
Innovation Solution
A user device selects a portion of media content using gestures, creates an image of the selected text, and sends it to a server for translation, which performs optical character recognition and language translation, returning the translated text and search results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple electronic translators are carried for different languages, then translation coverage is improved, but device complexity and portability are worsened
Solution Approach 1:
The patent implements a universal translation system that can translate multiple languages using a single device. The system includes a language identification module that detects the source language and a translation module that translates to the target language, enabling one device to perform the function of multiple language-specific translators.
Solution Approach 2:
The patent combines multiple translation functions into a single integrated system. Instead of using separate electronic translators for different language pairs, the system merges language detection, translation, and result display into one unified device that handles multiple languages.
2Loss of information
If camera captures entire scene, then context is preserved, but processing time and bandwidth are increased
Solution Approach 1:
The patent extracts only the relevant text portion from the captured image for translation. The system identifies and extracts the region containing text, creates a cropped image of just that portion, and sends only the extracted text or cropped image to the server, reducing processing time and bandwidth usage while preserving the necessary contextual information.
Solution Approach 2:
The patent segments the captured image to separate the text region from the rest of the scene. By dividing the image into relevant (text-containing) and irrelevant portions, the system processes only the necessary segment for translation, maintaining context while optimizing performance.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables fast and accurate translation of user-selected text with reduced network bandwidth usage and computational load on the device, providing quick and efficient language translation.
Implementation Method 1
the server to translate the one or more first characters. The server may perform optical character recognition on the received image to identify the one or more first characters
Data Source
AI summary
Some implementations disclosed herein provide techniques and arrangements to enable translating language characters in media content. For example, some implementations receive a user selection of a first portion of media content. Some implementations disclosed herein may, based on the first portion, identify a second portion of the media content. The second portion of the media content may include one or more first characters of a first language. Some implementations disclosed herein may create an image that includes the second portion of the media content and may send the image to a server. Some implementations disclosed herein may receive one or more second characters of a second language corresponding to a translation of the one or more first characters of the first language from the server.


