Augmented Reality Image Translation with Dynamic Frame Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image translation applications face limitations in translating moving or dynamically positioned objects, as they require manual capture and input, which can lead to misinterpretation and are not suitable for real-time translation in augmented reality scenarios.
Innovation Solution
A method and system for augmented reality-based image translation that captures video frames, extracts suitable frames, translates text in real-time, and renders the translation result onto the corresponding frames, allowing for continuous translation regardless of camera angle or object position changes, using a feature transformation model for accurate rendering.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a user manually captures and inputs text for translation, then translation accuracy can be maintained, but the process is time-consuming and prone to typos
Solution Approach 1:
The system enables automatic text extraction and translation without manual user input. The camera captures images, the recognition unit automatically identifies text, and the translation unit processes it, making the system serve itself rather than requiring manual operation for each step
Solution Approach 2:
The system performs preliminary actions by capturing multiple frames in advance and extracting text from them automatically. This allows the translation process to begin immediately without waiting for manual text input, reducing overall translation time while maintaining accuracy through automated recognition
2Reliability
If an image translation application captures static objects, then translation can be provided, but the application range is limited when objects move or the camera angle changes
Solution Approach 1:
The system transitions from static image processing to dynamic video frame processing. By continuously capturing multiple frames and tracking object positions across frames, the system adapts to moving objects and changing camera angles, maintaining translation service availability in dynamic environments
Solution Approach 2:
The system achieves universality by handling both static and dynamic scenarios through a unified approach. The same core functionality of capturing, recognizing, and translating text now works across diverse situations including moving objects, changing angles, and real-time video streams, significantly expanding application range
3Measurement precision
If text is translated manually or through static image capture, then translation results are accurate, but real-time translation in augmented reality scenarios cannot be achieved
Solution Approach 1:
The system enables continuous translation by processing video frames in real-time as they are captured. The recognition unit continuously identifies text in each frame, and the translation unit continuously translates it, providing uninterrupted real-time translation service that maintains accuracy through ongoing processing rather than batch processing
Data Source
AI summary
Provided is a method for augmented reality-based image translation performed by one or more processors, which includes storing a plurality of frames representing a video captured by a camera, extracting a first frame that satisfies a predetermined criterion from the stored plurality of frames, translating a first language sentence (or group of words) included in the first frame into a second language sentence (or group of words), determining a translation region including the second language sentence (or group of words) included in the first frame, and rendering the translation region in a second frame.


