Speech Translation Apparatus Automatic Speaker Direction Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech translation systems require users to perform button operations each time they speak, disrupting the conversation flow and reducing operability.
Innovation Solution
A speech translation apparatus that automatically switches between input and output languages based on the speaker's identity, determined by a sound source direction estimator and controller, using pre-stored layout information to display original and translated text on a user-friendly interface.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If automatic speaker identification is implemented using sound source direction estimation, then operability is improved by eliminating frequent button operations, but device complexity increases due to additional hardware and processing requirements
Solution Approach 1:
The system automatically identifies speakers and switches translation directions without user intervention. The sound source direction estimator and controller work autonomously to detect which user is speaking and adjust the translation output accordingly, eliminating the need for manual button operations and achieving self-service operation.
Solution Approach 2:
The patent replaces manual mechanical button operations with an automated acoustic field-based identification system. Instead of users physically pressing buttons to switch languages, the system uses microphone arrays and sound source direction estimation to automatically detect speaker identity and switch translation directions programmatically.
2Measurement precision
If sound source direction estimation is used to identify speakers, then accuracy in determining input language is improved, but measurement precision requirements increase due to noisy environments
Solution Approach 1:
The patent introduces a sound source direction estimator as an intermediary between the acoustic signals and the translation system. This intermediary component processes the raw acoustic signals from the microphone array, estimates sound source directions, and provides cleaned speaker identification information to the controller, filtering out noise interference in the process.
Solution Approach 2:
The system uses multiple microphones in an array to create multiple copies of the acoustic signal from different spatial perspectives. By comparing these copied signals and their directional information, the system can accurately identify the sound source direction even in noisy environments, as the spatial redundancy helps filter out unrelated noise.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enhances operability by allowing seamless conversation without the need for frequent button operations, improving natural interaction and accuracy in language recognition and translation, even in noisy environments.
Implementation Method 1
a sound source direction estimator which estimates a sound source direction by processing an acoustic signal obtained by a microphone array unit
Data Source
AI summary
A speech translation apparatus includes: an estimator which estimates a sound source direction, based on an acoustic signal obtained by a microphone array unit; a controller which identifies that an utterer is a user or a conversation partner, based on the sound source direction estimated after the start of translation is instructed by a button, using a positional relationship indicated by a layout information item stored in storage and selected in advance, and determines a translation direction indicating input and output languages in and into which content of the acoustic signal is recognized and translated, respectively; and a translator which obtains, according to the translation direction, original text indicating the content in the input language and translated text indicating the content in the output language. The controller displays the original and translated texts on first and second display areas corresponding to the positions of the user and conversation partner, respectively.


