Augmented Reality Image Translation with Dynamic Frame Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image translation applications face limitations in translating moving or dynamically positioned objects, as they require manual capture and input, which can lead to misinterpretation and are not suitable for real-time translation in augmented reality scenarios.

Innovation Solution

A method and system for augmented reality-based image translation that captures video frames, extracts suitable frames, translates text in real-time, and renders the translation result onto the corresponding frames, allowing for continuous translation regardless of camera angle or object position changes, using a feature transformation model for accurate rendering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a user manually captures and inputs text for translation, then translation accuracy can be maintained, but the process is time-consuming and prone to typos

Engineering Contradiction:
Improvetranslation accuracyVSAvoidtranslation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables automatic text extraction and translation without manual user input. The camera captures images, the recognition unit automatically identifies text, and the translation unit processes it, making the system serve itself rather than requiring manual operation for each step

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary actions by capturing multiple frames in advance and extracting text from them automatically. This allows the translation process to begin immediately without waiting for manual text input, reducing overall translation time while maintaining accuracy through automated recognition

Inventive Principle:
Principle #10Preliminary action

2Reliability

If an image translation application captures static objects, then translation can be provided, but the application range is limited when objects move or the camera angle changes

Engineering Contradiction:
Improvetranslation service availabilityVSAvoidapplication range
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system transitions from static image processing to dynamic video frame processing. By continuously capturing multiple frames and tracking object positions across frames, the system adapts to moving objects and changing camera angles, maintaining translation service availability in dynamic environments

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system achieves universality by handling both static and dynamic scenarios through a unified approach. The same core functionality of capturing, recognizing, and translating text now works across diverse situations including moving objects, changing angles, and real-time video streams, significantly expanding application range

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If text is translated manually or through static image capture, then translation results are accurate, but real-time translation in augmented reality scenarios cannot be achieved

Engineering Contradiction:
Improvetranslation accuracyVSAvoidtranslation speed
Core Design Contradiction:
Measurement precisionVSSpeed

Solution Approach 1:

The system enables continuous translation by processing video frames in real-time as they are captured. The recognition unit continuously identifies text in each frame, and the translation unit continuously translates it, providing uninterrupted real-time translation service that maintains accuracy through ongoing processing rather than batch processing

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS12197883B2Method and system for image translation
Publication Date: 2025.01.14 NAVER CORP
  • US12197883B2 patent drawing
  • US12197883B2 patent drawing
  • US12197883B2 patent drawing

AI summary

Provided is a method for augmented reality-based image translation performed by one or more processors, which includes storing a plurality of frames representing a video captured by a camera, extracting a first frame that satisfies a predetermined criterion from the stored plurality of frames, translating a first language sentence (or group of words) included in the first frame into a second language sentence (or group of words), determining a translation region including the second language sentence (or group of words) included in the first frame, and rendering the translation region in a second frame.