AR Glasses Text Translation via Image Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Augmented reality (AR) glasses lack a translation function for text appearing in a user's vicinity, hindering user experience by not being able to translate and display text from the environment in real-time.
Innovation Solution
AR glasses collect environment images and, upon a user's trigger instruction, use image recognition and AI to identify text, obtain the translation result, and display it in a preset mode within a virtual AR space constructed based on the real environment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If AR glasses are used to view the real environment, then users can see virtual objects superimposed on the physical environment, but the glasses lack text translation function which hinders user experience
Solution Approach 1:
The AR glasses are enhanced with multiple functions including text recognition, translation, and display capabilities. The processing unit not only renders virtual images but also recognizes text from captured images, translates it using translation data, and displays both original and translated text in the AR space, making the device universal for viewing, translating, and interacting with text in the environment.
Solution Approach 2:
The processing unit acts as an intermediary between the camera capture and the display output. It receives captured images, extracts text, translates the text using translation data, and then displays the translated text in the AR space, mediating the transformation from raw visual input to meaningful translated output for the user.
2Productivity
If text translation and display functions are added to AR glasses, then real-time translation is achieved, but device complexity increases
Solution Approach 1:
The text recognition, translation, and display functions are merged into the existing AR glasses architecture. The processing unit combines image capture, text extraction, translation processing, and AR display operations into a unified system, where the same hardware components serve multiple functions, thereby achieving real-time translation without proportionally increasing device complexity.
Solution Approach 2:
The processing unit is designed to perform multiple functions: rendering virtual images, recognizing text from captured images, translating recognized text, and displaying both virtual objects and translated text in the AR space. This multi-functionality allows the device to achieve real-time translation capability while utilizing existing hardware resources efficiently.
Data Source
AI summary
A translation method is performed by an augmented reality (AR) device. The method includes obtaining a text to be translated from an environment image in response to a translation trigger instruction; obtaining a translation result of the text to be translated; and displaying the translation result based on a preset display mode in an AR space. The AR space is a virtual space constructed by the AR device based on a real environment.

