Low Latency Multi-Language Translation via Local Wireless Networking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Real-time translation in multi-language settings is hindered by the need for online connectivity and the complexity of selecting and switching between languages, especially in situations where groups of people speak multiple languages and online connectivity is limited or non-existent.

Innovation Solution

Establishing a multi-language translation group through local wireless networking, allowing computing devices to perform automated speech recognition and translation, generating non-audio data that can be communicated among users, and converting it back to spoken audio in their desired languages, using technologies like Bluetooth, NFC, or WLAN for communication.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If online translation service is used, then translation functionality is provided, but online connectivity is required and latency increases

Engineering Contradiction:
Improvetranslation functionalityVSAvoidlatency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The translation system is segmented into multiple distributed computing devices, each capable of performing speech recognition, translation, and text-to-speech conversion independently. This distribution eliminates the need for centralized online processing, reducing latency and enabling offline operation while maintaining translation functionality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A local wireless network acts as an intermediary between computing devices, enabling direct peer-to-peer communication for translation operations without requiring external online services. This intermediary network layer allows devices to exchange speech data and translated output locally, reducing dependency on online connectivity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If explicit language designation is required, then translation accuracy is improved, but user operation complexity increases

Engineering Contradiction:
Improvetranslation accuracyVSAvoidlanguage selection complexity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system automatically detects and identifies the language being spoken by analyzing speech patterns and metadata, eliminating the need for users to manually select languages. Each computing device self-determines the source and target languages based on the conversation context, maintaining translation accuracy while simplifying user interaction.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system continuously monitors speech inputs and provides feedback about detected languages to users, allowing automatic language selection while giving users the option to override if needed. This feedback mechanism ensures accurate language identification without requiring explicit user designation.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If local wireless networking is used, then online connectivity dependency is reduced, but device complexity increases

Engineering Contradiction:
Improveoffline capabilityVSAvoidnetworking complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

Standard computing devices equipped with common wireless networking capabilities (Bluetooth, Wi-Fi Direct, NFC) are used for translation operations. These ubiquitous technologies enable offline translation functionality without requiring specialized hardware or complex custom networking solutions, maintaining device simplicity while achieving versatility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Device complexity

If translation is performed by a single device, then processing is simplified, but processing overhead and latency increase

Engineering Contradiction:
Improveprocessing simplicityVSAvoidtranslation speed
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The translation workload is segmented and distributed across multiple computing devices in the group. Each device processes speech inputs independently through speech recognition and translation, then shares results with other devices. This parallel processing approach reduces individual device overhead and decreases overall translation latency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Multiple computing devices are merged into a collaborative translation network where each device contributes processing capabilities. The combined computational power of all devices in the group handles translation operations more efficiently than a single device, improving translation speed while distributing processing complexity across the network.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP3602545B1Low latency nearby group translation
Publication Date: 2021.11.24 GOOGLE LLC
  • EP3602545B1 patent drawingFigure 1
  • EP3602545B1 patent drawingFigure 2A
  • EP3602545B1 patent drawingFigure 2B

AI summary

Local wireless networking is used to establish a multi-language translation group between users that speak more than two different languages such that when one user speaks into his or her computing device, that user's computing device may perform automated speech recognition, and in some instances, translation to a different language, to generate non-audio data for communication to the computing devices of other users in the group. The other users' computing devices then generate spoken audio outputs suitable for their respective users using the received non-audio data. The generation of the spoken audio outputs on the other users' computing devices may also include performing a translation, thereby enabling each user to receive spoken audio output in their desired language in response to speech input from another user, and irrespective of the original language of the speech input.