Mobile Terminal Voice Source Separation for Multi-Speaker Translation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In environments with multiple speakers, existing technologies fail to effectively separate voice signals from individual speakers, leading to mixed voices and difficulties in language translation.
Innovation Solution
A mobile terminal equipped with a microphone to generate voice signals, a processor for voice source separation based on speaker positions, and memory to store language information, allowing for the generation and translation of separated voice signals into target languages.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If a microphone receives all voices from multiple speakers in a meeting room or classroom, then the microphone can capture all spoken content, but the voice signals become mixed and cannot be separated by speaker
Solution Approach 1:
The patent segments the mixed voice signal into separate voice signals for each speaker by performing voice source separation based on voice source positions. The processor divides the composite voice signal captured by the microphone into individual speaker signals using spatial information from multiple microphones to identify and separate each speaker's contribution.
Solution Approach 2:
The patent introduces spatial dimension (voice source position) as an additional parameter to separate mixed voice signals. By utilizing position information from multiple microphones arranged in space, the system separates speakers not just by time or frequency but by their spatial locations, adding a dimensional aspect to the signal separation process.
2Measurement precision
If voice source separation is performed based on voice source positions, then separated voice signals for specific speakers can be generated, but the device complexity increases due to multiple microphones and processing requirements
Solution Approach 1:
The patent makes the mobile terminal perform multiple functions: it acts as both a communication device and a voice separation/translation system. The processor handles both standard mobile operations and the complex task of voice source separation and real-time translation, reducing the need for dedicated separate hardware systems.
Solution Approach 2:
The mobile terminal uses its own built-in microphones and processor to perform voice source separation and translation without requiring external specialized equipment. The system processes and translates voices independently using its own resources, making the complex functionality self-contained within the mobile device.
3Reliability
If translation results are generated for separated voice signals, then accurate translation can be provided, but the processing time and computational load increase
Solution Approach 1:
The patent performs voice source separation before translation, preparing the separated voice signals in advance. By separating the mixed signals first and organizing them by speaker position, the translation process receives pre-processed, organized input, which streamlines the subsequent translation operation and reduces overall processing time.
Data Source
AI summary
A mobile terminal is disclosed. The mobile terminal comprises: a microphone configured to generate a voice signal in response to voices of speakers; a processor configured to generate a separated voice signal associated with each of the voices by separating the voice signal from a sound source on the basis of a sound source location of each of the voices, and output the result of translation for each of the voices, on the basis of the separated voice signal; and a memory configured to store source language information indicating source languages that are uttered languages of the voices of the speakers. The processor outputs the results of translations in which the languages of the voices of the speakers have been translated from the source languages into a target language, on the basis of the source language information and the separated voice signal.


