Real-Time Voice Translation Echo Cancellation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Electronic devices struggle to accurately translate voices in real-time conversations when connected to wireless input/output devices, often resulting in overlapping voices and increased waiting times due to echo issues and noise interference.

Innovation Solution

An electronic device equipped with microphones, speakers, and a processor that uses echo cancellation and machine learning to separate and translate voices in real-time, even when voices overlap, by preprocessing audio data to distinguish between user and counterpart voices based on volume and noise cancellation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If the electronic device translates and outputs the counterpart's voice through the wireless input/output device, then the translation service is provided, but the user's voice and the translated voice overlap causing translation failure

Engineering Contradiction:
Improvetranslation service availabilityVSAvoidtranslation accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent divides the audio processing into separate channels: the electronic device processes the user's voice while the wireless input/output device processes the counterpart's voice. This segmentation prevents voice overlap by assigning different translation tasks to different devices, thereby maintaining both translation service availability and accuracy.

Inventive Principle:
Principle #1Segmentation

2Reliability

If the electronic device waits for the user to finish speaking before translating, then translation accuracy is maintained, but waiting time increases

Engineering Contradiction:
Improvetranslation accuracyVSAvoidwaiting time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements preliminary voice separation and buffering before translation. The electronic device captures and buffers the user's voice in advance, performs voice activity detection and separation, and prepares translation data structures beforehand. This preliminary processing enables the system to maintain high translation accuracy while significantly reducing the perceived waiting time by having translation ready for immediate output.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If the electronic device processes both user and counterpart voices simultaneously, then real-time translation is achieved, but voice separation becomes difficult

Engineering Contradiction:
Improvetranslation speedVSAvoidvoice separation difficulty
Core Design Contradiction:
ProductivityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent introduces an intermediary wireless input/output device that acts as a mediator in the translation process. This device captures the counterpart's voice directly, performing voice activity detection and separation locally. By using this intermediary, the electronic device can process both voices simultaneously without direct interference, as the intermediary handles the counterpart's voice processing and returns only the separated translation data, thereby maintaining real-time translation speed while simplifying voice separation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables accurate and simultaneous translation of user and counterpart voices during conversations, improving user convenience by effectively canceling noise and echo, even when voices overlap, and reducing waiting times.

Implementation Method 1

generate first audio data by cancelling an echo from the acquired first audio

Methodology Applied
Scientific EffectEcho cancellation: Echo

Data Source

PatentUS20240020490A1Method and apparatus for processing translation
Publication Date: 2024.01.18 SAMSUNG ELECTRONICS CO LTD
  • US20240020490A1 patent drawing
  • US20240020490A1 patent drawing
  • US20240020490A1 patent drawing

AI summary

An electronic device includes a processor configured to: acquire first audio a microphone being connected with an external device through a communication module; generate first audio data by cancelling an echo from the acquired first audio; transmit the first audio data to an external device; receive at least one of second audio or second audio data, the second audio or the second audio data being acquired through a microphone of the external device from the external device; translate the first audio data and obtain first translation information; translate the second audio data and obtain second translation information; transmit the first translation information to the external device; and output the second translation information.