Speech Translation Apparatus Automatic Speaker Direction Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech translation systems require users to perform button operations each time they speak, disrupting the conversation flow and reducing operability.

Innovation Solution

A speech translation apparatus that automatically switches between input and output languages based on the speaker's identity, determined by a sound source direction estimator and controller, using pre-stored layout information to display original and translated text on a user-friendly interface.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If automatic speaker identification is implemented using sound source direction estimation, then operability is improved by eliminating frequent button operations, but device complexity increases due to additional hardware and processing requirements

Engineering Contradiction:
ImproveoperabilityVSAvoiddevice complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system automatically identifies speakers and switches translation directions without user intervention. The sound source direction estimator and controller work autonomously to detect which user is speaking and adjust the translation output accordingly, eliminating the need for manual button operations and achieving self-service operation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical button operations with an automated acoustic field-based identification system. Instead of users physically pressing buttons to switch languages, the system uses microphone arrays and sound source direction estimation to automatically detect speaker identity and switch translation directions programmatically.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If sound source direction estimation is used to identify speakers, then accuracy in determining input language is improved, but measurement precision requirements increase due to noisy environments

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidnoisy environment interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent introduces a sound source direction estimator as an intermediary between the acoustic signals and the translation system. This intermediary component processes the raw acoustic signals from the microphone array, estimates sound source directions, and provides cleaned speaker identification information to the controller, filtering out noise interference in the process.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system uses multiple microphones in an array to create multiple copies of the acoustic signal from different spatial perspectives. By comparing these copied signals and their directional information, the system can accurately identify the sound source direction even in noisy environments, as the spatial redundancy helps filter out unrelated noise.

Inventive Principle:
Principle #26Copying

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enhances operability by allowing seamless conversation without the need for frequent button operations, improving natural interaction and accuracy in language recognition and translation, even in noisy environments.

Implementation Method 1

a sound source direction estimator which estimates a sound source direction by processing an acoustic signal obtained by a microphone array unit

Methodology Applied
Scientific EffectAcoustic signal processing: Acoustics

Data Source

PatentUS11182567B2Speech translation apparatus, speech translation method, and recording medium storing the speech translation method
Publication Date: 2021.11.23 PANASONIC HOLDINGS CORP
  • US11182567B2 patent drawing
  • US11182567B2 patent drawing
  • US11182567B2 patent drawing

AI summary

A speech translation apparatus includes: an estimator which estimates a sound source direction, based on an acoustic signal obtained by a microphone array unit; a controller which identifies that an utterer is a user or a conversation partner, based on the sound source direction estimated after the start of translation is instructed by a button, using a positional relationship indicated by a layout information item stored in storage and selected in advance, and determines a translation direction indicating input and output languages in and into which content of the acoustic signal is recognized and translated, respectively; and a translator which obtains, according to the translation direction, original text indicating the content in the input language and translated text indicating the content in the output language. The controller displays the original and translated texts on first and second display areas corresponding to the positions of the user and conversation partner, respectively.