AR Glasses Text Translation via Image Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Augmented reality (AR) glasses lack a translation function for text appearing in a user's vicinity, hindering user experience by not being able to translate and display text from the environment in real-time.

Innovation Solution

AR glasses collect environment images and, upon a user's trigger instruction, use image recognition and AI to identify text, obtain the translation result, and display it in a preset mode within a virtual AR space constructed based on the real environment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If AR glasses are used to view the real environment, then users can see virtual objects superimposed on the physical environment, but the glasses lack text translation function which hinders user experience

Engineering Contradiction:
Improvetranslation functionVSAvoiduser experience
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The AR glasses are enhanced with multiple functions including text recognition, translation, and display capabilities. The processing unit not only renders virtual images but also recognizes text from captured images, translates it using translation data, and displays both original and translated text in the AR space, making the device universal for viewing, translating, and interacting with text in the environment.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The processing unit acts as an intermediary between the camera capture and the display output. It receives captured images, extracts text, translates the text using translation data, and then displays the translated text in the AR space, mediating the transformation from raw visual input to meaningful translated output for the user.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If text translation and display functions are added to AR glasses, then real-time translation is achieved, but device complexity increases

Engineering Contradiction:
Improvereal-time translation capabilityVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The text recognition, translation, and display functions are merged into the existing AR glasses architecture. The processing unit combines image capture, text extraction, translation processing, and AR display operations into a unified system, where the same hardware components serve multiple functions, thereby achieving real-time translation without proportionally increasing device complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The processing unit is designed to perform multiple functions: rendering virtual images, recognizing text from captured images, translating recognized text, and displaying both virtual objects and translated text in the AR space. This multi-functionality allows the device to achieve real-time translation capability while utilizing existing hardware resources efficiently.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12008696B2Translation method and AR device
Publication Date: 2024.06.11 BEIJING XIAOMI MOBILE SOFTWARE CO LTD
  • US12008696B2 patent drawing
  • US12008696B2 patent drawing

AI summary

A translation method is performed by an augmented reality (AR) device. The method includes obtaining a text to be translated from an environment image in response to a translation trigger instruction; obtaining a translation result of the text to be translated; and displaying the translation result based on a preset display mode in an AR space. The AR space is a virtual space constructed by the AR device based on a real environment.