On-Device Bilingual Translation Model for AR Headsets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional machine translation devices often require remote server connections for translation, which can be inefficient and lack the ability to understand context and nuances in text, especially when dealing with multiple languages and formats.

Innovation Solution

A head-mounted display equipped with a machine translation model trained on various tasks and datasets, including multiple language directions, to translate text in real-time, using image sensors to capture text and perform OCR, and ASR to handle spoken language, with the capability to format and normalize text accurately.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If machine translation is performed using remote server connections, then translation capability is provided, but translation efficiency and speed deteriorate due to network dependency and latency

Engineering Contradiction:
Improvetranslation capabilityVSAvoidtranslation efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent extracts the machine translation model from remote servers and embeds it directly into the head-mounted display device. This allows the translation functionality to operate locally without network dependency, eliminating latency and improving translation speed while maintaining reliability through self-contained processing capability

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The head-mounted display performs translation operations autonomously using its own embedded machine translation model, image sensor for text capture, and processing unit. This self-service approach eliminates the need for external server connections, enabling real-time translation with immediate response to captured text

Inventive Principle:
Principle #25Self-service

2Ease of manufacture

If conventional machine translation models are used, then basic translation is provided, but understanding of context and nuances deteriorates

Engineering Contradiction:
Improvetranslation functionalityVSAvoidcontext understanding
Core Design Contradiction:
Ease of manufactureVSLoss of information

Solution Approach 1:

The machine translation model undergoes extensive pre-training on diverse datasets including multiple languages, text formats, and contextual scenarios before deployment in the head-mounted display. This preliminary training equips the model with enhanced context understanding and nuance recognition capabilities, allowing it to accurately interpret and translate text while preserving meaning across different languages and formats

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent employs advanced model training techniques that adjust key parameters such as attention mechanisms, embedding dimensions, and training data composition. These parameter changes enable the translation model to better capture contextual relationships and linguistic nuances, significantly improving translation quality compared to conventional models

Inventive Principle:
Principle #35Parameter changes

3Productivity

If text translation is performed in real-time, then translation speed is improved, but processing demands and memory requirements increase

Engineering Contradiction:
Improvetranslation speedVSAvoidmemory and processing demands
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The translation processing pipeline is segmented into distinct stages: image capture by the image sensor, OCR text recognition, machine translation processing, and result display. This segmentation allows each component to operate independently and efficiently, enabling real-time translation while optimizing resource utilization and reducing peak processing demands on the device

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system processes text in partial segments rather than complete documents at once. The image sensor captures text incrementally, the OCR processes visible portions, and the translation model translates segments as they become available. This partial processing approach enables real-time translation output while keeping memory and processing requirements at manageable levels

Inventive Principle:
Principle #16Partial or excessive action

4Adaptability or versatility

If multiple languages and text formats are handled, then translation versatility is improved, but model complexity and training requirements increase

Engineering Contradiction:
Improvemulti-language supportVSAvoidmodel complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The machine translation model is designed as a universal system capable of handling multiple languages and diverse text formats simultaneously. The model architecture incorporates multi-language training data and format-agnostic processing capabilities, allowing a single model to perform translation across numerous language pairs and text types without requiring separate specialized models for each language or format

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent extends the translation model's capability by adding dimensional capacity to process multiple languages and formats within the same model structure. This is achieved through enhanced embedding spaces that accommodate diverse linguistic features and training approaches that incorporate varied text formats, enabling the model to handle complex multi-language scenarios without proportionally increasing overall system complexity

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20250103831A1Bilingual multitask machine translation model for live translation on artificial reality devices
Publication Date: 2025.03.27 META PLATFORMS INC
  • US20250103831A1 patent drawing
  • US20250103831A1 patent drawing
  • US20250103831A1 patent drawing

AI summary

Head-mounted displays may include a machine translation model designed to recognize text through optical character recognition or automatic speech recognition, and may translate the text from its original language to another language. The machine translation model may be trained to modify source text using various tasks, thus allowing the machine translation model to learn different versions of the source text in several different versions. The source text and a variation(s) derived from a task(s) may be mapped to a target text, representing the properly translated and formatted version of the source text. The machine translation model may provide a single model, to facilitate machine translation, implemented on the head-mounted display. Also, the machine translation model may include a bilingual machine translation model that may translate source text from one language to another language, and vice versa.