Speaker Separation for Multi-Speaker Interpretation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automatic interpretation systems are limited to face-to-face conversations and fail to effectively interpret multiple speakers' speech in surrounding environments, such as foreign languages heard during travel or daily interactions.
Innovation Solution
A system and method that separates and interprets multiple speech signals from a user and their surroundings, converting them into the desired language, using a user terminal with a communication module, processor, and memory to perform speaker-specific speech separation, interpretation, and result classification, reflecting intensity and echo information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional automatic interpretation systems are used for face-to-face conversations, then interpretation accuracy is improved, but the system cannot interpret surrounding speech from multiple speakers
Solution Approach 1:
The patent segments the mixed speech signal into separate speaker-specific signals using speaker separation technology. This allows the system to identify and process individual speakers' speech separately, enabling accurate interpretation of surrounding speech from multiple speakers while maintaining face-to-face conversation capabilities.
Solution Approach 2:
The patent creates a universal interpretation system that can handle multiple speech scenarios including face-to-face conversations, surrounding speech from multiple speakers, and mixed environments. The system uses speaker separation to identify the source of speech and applies appropriate interpretation methods for each scenario, making it adaptable to diverse communication contexts.
2Adaptability or versatility
If speaker separation is performed on mixed speech signals, then interpretation of multiple speakers is enabled, but system complexity increases
Solution Approach 1:
The patent introduces speaker separation technology as an intermediary processing step between the microphone and the interpretation module. This intermediary component analyzes the mixed speech signal, identifies individual speakers based on acoustic characteristics, and routes each speaker's speech to the appropriate interpretation handler, simplifying the overall system architecture while enabling multi-speaker capability.
3Loss of information
If all speech signals are interpreted into the user's language, then comprehensive information is provided, but difficulty in identifying which speaker said what increases
Solution Approach 1:
The patent applies local quality by providing differentiated information for each speaker's speech. Instead of uniformly processing all speech, the system identifies and processes speech from different speakers separately, attaching speaker identification metadata to each interpreted speech segment. This allows users to understand both the content and the source of each speech act.
Solution Approach 2:
The patent implements feedback mechanisms where the system continuously monitors the speech environment, identifies speakers based on acoustic fingerprints and spatial information, and provides real-time feedback about which speaker is speaking. This feedback loop helps users track the source of speech while maintaining comprehensive interpretation of all speakers.
Data Source
AI summary
Provided is a method of performing automatic interpretation based on speaker separation by a user terminal, the method including: receiving a first speech signal including at least one of a user speech of a user and a user surrounding speech around the user from an automatic interpretation service providing terminal, separating the first speech signal into speaker-specific speech signals, performing interpretation on the speaker-specific speech signals in a language selected by the user on the basis of an interpretation mode, and providing a second speech signal generated as a result of the interpretation to at least one of a counterpart terminal and the automatic interpretation service providing terminal according to the interpretation mode.


