Voice Processing Apparatus Sound Source Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In environments with multiple speakers, existing technologies fail to effectively separate and output individual voice signals, leading to mixed voices and increased noise interference, making it difficult to selectively listen to or watch conversations based on speaker importance.

Innovation Solution

A voice processing apparatus and method that uses a microphone, communication circuit, memory, and processor to generate separation voice signals by determining sound source positions and outputting them according to set modes, allowing for selective listening or watching of individual speakers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If a single microphone receives all voices from multiple speakers, then the microphone can capture all speech signals, but the voices become mixed and cannot be separated by speaker

Engineering Contradiction:
Improvevoice signal captureVSAvoidspeaker identity information
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent segments the mixed voice signal into separate speaker-specific signals by determining sound source positions and generating separation voice signals for each speaker. The processor divides the captured audio stream into distinct components associated with different spatial locations, enabling individual speaker identification and selective output.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces sound source position information as an intermediary parameter to separate mixed voices. By using spatial position data as a mediating factor, the system can distinguish between different speakers in the mixed signal and generate separate output channels for each speaker based on their determined positions.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If all speaker voices are output simultaneously, then all speech information is conveyed, but noise interference increases and selective listening is not possible

Engineering Contradiction:
Improvespeech information completenessVSAvoidnoise interference
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The patent implements dynamic output control where the system can adaptively select which speaker signals to output based on determined sound source positions and user needs. The output mode can dynamically switch between presenting all speakers, selected speakers, or individual speakers, allowing flexible noise management while preserving speech information completeness.

Inventive Principle:
Principle #15Dynamics

3Loss of information

If voice separation is implemented based on sound source position, then individual speaker voices can be separated, but the system complexity increases

Engineering Contradiction:
Improvespeaker identity informationVSAvoidsignal processing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent employs self-service mechanisms where the system automatically determines sound source positions and generates separation voice signals without requiring manual intervention. The processor autonomously performs position determination, signal separation, and output mode selection based on the captured audio and stored position information, reducing the need for complex external control systems.

Inventive Principle:
Principle #25Self-service

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The solution minimizes noise interference and allows users to selectively listen to or watch specific speakers by generating separation voice signals based on sound source positions, enabling clearer communication in multi-speaker environments.

Implementation Method 1

a microphone configured to generate voice signals in response to voices of the plurality of speakers

Methodology Applied
Scientific EffectElectroacoustic transduction:

Data Source

PatentUS20240257824A1Voice processing apparatus for processing voices, voice processing system, and voice processing method
Publication Date: 2024.08.01 AMOSENSE CO LTD
  • US20240257824A1 patent drawing
  • US20240257824A1 patent drawing
  • US20240257824A1 patent drawing

AI summary

Disclosed is a voice processing apparatus for processing voices of a plurality of speakers. The voice processing apparatus comprises: a microphone configured to generate voice signals in response to the voices of the plurality of speakers; a communication circuit configured to transmit and receive data; memory; and a processor, wherein the processor, on the basis of instructions stored in the memory, performs sound source separation of the voice signals on the basis of sound source positions of each of the voices, generates separate voice signals associated with each of the voices according to the sound source separation, determines output modes corresponding to the sound source positions of each of the voices, and uses the communication circuit to output the separate voice signals according to the determined output modes.