Speaker Separation for Multi-Speaker Interpretation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automatic interpretation systems are limited to face-to-face conversations and fail to effectively interpret multiple speakers' speech in surrounding environments, such as foreign languages heard during travel or daily interactions.

Innovation Solution

A system and method that separates and interprets multiple speech signals from a user and their surroundings, converting them into the desired language, using a user terminal with a communication module, processor, and memory to perform speaker-specific speech separation, interpretation, and result classification, reflecting intensity and echo information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional automatic interpretation systems are used for face-to-face conversations, then interpretation accuracy is improved, but the system cannot interpret surrounding speech from multiple speakers

Engineering Contradiction:
Improveinterpretation scopeVSAvoidinterpretation accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the mixed speech signal into separate speaker-specific signals using speaker separation technology. This allows the system to identify and process individual speakers' speech separately, enabling accurate interpretation of surrounding speech from multiple speakers while maintaining face-to-face conversation capabilities.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal interpretation system that can handle multiple speech scenarios including face-to-face conversations, surrounding speech from multiple speakers, and mixed environments. The system uses speaker separation to identify the source of speech and applies appropriate interpretation methods for each scenario, making it adaptable to diverse communication contexts.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If speaker separation is performed on mixed speech signals, then interpretation of multiple speakers is enabled, but system complexity increases

Engineering Contradiction:
Improvemulti-speaker interpretation capabilityVSAvoidsignal processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces speaker separation technology as an intermediary processing step between the microphone and the interpretation module. This intermediary component analyzes the mixed speech signal, identifies individual speakers based on acoustic characteristics, and routes each speaker's speech to the appropriate interpretation handler, simplifying the overall system architecture while enabling multi-speaker capability.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If all speech signals are interpreted into the user's language, then comprehensive information is provided, but difficulty in identifying which speaker said what increases

Engineering Contradiction:
Improveinformation completenessVSAvoidspeaker identification difficulty
Core Design Contradiction:
Loss of informationVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies local quality by providing differentiated information for each speaker's speech. Instead of uniformly processing all speech, the system identifies and processes speech from different speakers separately, attaching speaker identification metadata to each interpreted speech segment. This allows users to understand both the content and the source of each speech act.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements feedback mechanisms where the system continuously monitors the speech environment, identifies speakers based on acoustic fingerprints and spatial information, and provides real-time feedback about which speaker is speaking. This feedback loop helps users track the source of speech while maintaining comprehensive interpretation of all speakers.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12112769B2System, user terminal, and method for providing automatic interpretation service based on speaker separation
Publication Date: 2024.10.08 ELECTRONICS & TELECOMM RES INST
  • US12112769B2 patent drawing
  • US12112769B2 patent drawing
  • US12112769B2 patent drawing

AI summary

Provided is a method of performing automatic interpretation based on speaker separation by a user terminal, the method including: receiving a first speech signal including at least one of a user speech of a user and a user surrounding speech around the user from an automatic interpretation service providing terminal, separating the first speech signal into speaker-specific speech signals, performing interpretation on the speaker-specific speech signals in a language selected by the user on the basis of an interpretation mode, and providing a second speech signal generated as a result of the interpretation to at least one of a counterpart terminal and the automatic interpretation service providing terminal according to the interpretation mode.