Mobile Terminal Voice Source Separation for Multi-Speaker Translation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In environments with multiple speakers, existing technologies fail to effectively separate voice signals from individual speakers, leading to mixed voices and difficulties in language translation.

Innovation Solution

A mobile terminal equipped with a microphone to generate voice signals, a processor for voice source separation based on speaker positions, and memory to store language information, allowing for the generation and translation of separated voice signals into target languages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If a microphone receives all voices from multiple speakers in a meeting room or classroom, then the microphone can capture all spoken content, but the voice signals become mixed and cannot be separated by speaker

Engineering Contradiction:
Improvevoice signal captureVSAvoidspeaker identity information
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent segments the mixed voice signal into separate voice signals for each speaker by performing voice source separation based on voice source positions. The processor divides the composite voice signal captured by the microphone into individual speaker signals using spatial information from multiple microphones to identify and separate each speaker's contribution.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces spatial dimension (voice source position) as an additional parameter to separate mixed voice signals. By utilizing position information from multiple microphones arranged in space, the system separates speakers not just by time or frequency but by their spatial locations, adding a dimensional aspect to the signal separation process.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If voice source separation is performed based on voice source positions, then separated voice signals for specific speakers can be generated, but the device complexity increases due to multiple microphones and processing requirements

Engineering Contradiction:
Improvevoice source position accuracyVSAvoidmicrophone array and processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the mobile terminal perform multiple functions: it acts as both a communication device and a voice separation/translation system. The processor handles both standard mobile operations and the complex task of voice source separation and real-time translation, reducing the need for dedicated separate hardware systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The mobile terminal uses its own built-in microphones and processor to perform voice source separation and translation without requiring external specialized equipment. The system processes and translates voices independently using its own resources, making the complex functionality self-contained within the mobile device.

Inventive Principle:
Principle #25Self-service

3Reliability

If translation results are generated for separated voice signals, then accurate translation can be provided, but the processing time and computational load increase

Engineering Contradiction:
Improvetranslation accuracyVSAvoidtranslation processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs voice source separation before translation, preparing the separated voice signals in advance. By separating the mixed signals first and organizing them by speaker position, the translation process receives pre-processed, organized input, which streamlines the subsequent translation operation and reduces overall processing time.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20230377594A1Mobile terminal capable of processing voice and operation method therefor
Publication Date: 2023.11.23 AMOSENSE CO LTD
  • US20230377594A1 patent drawing
  • US20230377594A1 patent drawing
  • US20230377594A1 patent drawing

AI summary

A mobile terminal is disclosed. The mobile terminal comprises: a microphone configured to generate a voice signal in response to voices of speakers; a processor configured to generate a separated voice signal associated with each of the voices by separating the voice signal from a sound source on the basis of a sound source location of each of the voices, and output the result of translation for each of the voices, on the basis of the separated voice signal; and a memory configured to store source language information indicating source languages that are uttered languages of the voices of the speakers. The processor outputs the results of translations in which the languages of the voices of the speakers have been translated from the source languages into a target language, on the basis of the source language information and the separated voice signal.