Voice Processing Apparatus Sound Source Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In environments with multiple speakers, existing technologies fail to effectively separate and output individual voice signals, leading to mixed voices and increased noise interference, making it difficult to selectively listen to or watch conversations based on speaker importance.
Innovation Solution
A voice processing apparatus and method that uses a microphone, communication circuit, memory, and processor to generate separation voice signals by determining sound source positions and outputting them according to set modes, allowing for selective listening or watching of individual speakers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If a single microphone receives all voices from multiple speakers, then the microphone can capture all speech signals, but the voices become mixed and cannot be separated by speaker
Solution Approach 1:
The patent segments the mixed voice signal into separate speaker-specific signals by determining sound source positions and generating separation voice signals for each speaker. The processor divides the captured audio stream into distinct components associated with different spatial locations, enabling individual speaker identification and selective output.
Solution Approach 2:
The patent introduces sound source position information as an intermediary parameter to separate mixed voices. By using spatial position data as a mediating factor, the system can distinguish between different speakers in the mixed signal and generate separate output channels for each speaker based on their determined positions.
2Loss of information
If all speaker voices are output simultaneously, then all speech information is conveyed, but noise interference increases and selective listening is not possible
Solution Approach 1:
The patent implements dynamic output control where the system can adaptively select which speaker signals to output based on determined sound source positions and user needs. The output mode can dynamically switch between presenting all speakers, selected speakers, or individual speakers, allowing flexible noise management while preserving speech information completeness.
3Loss of information
If voice separation is implemented based on sound source position, then individual speaker voices can be separated, but the system complexity increases
Solution Approach 1:
The patent employs self-service mechanisms where the system automatically determines sound source positions and generates separation voice signals without requiring manual intervention. The processor autonomously performs position determination, signal separation, and output mode selection based on the captured audio and stored position information, reducing the need for complex external control systems.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The solution minimizes noise interference and allows users to selectively listen to or watch specific speakers by generating separation voice signals based on sound source positions, enabling clearer communication in multi-speaker environments.
Implementation Method 1
a microphone configured to generate voice signals in response to voices of the plurality of speakers
Data Source
AI summary
Disclosed is a voice processing apparatus for processing voices of a plurality of speakers. The voice processing apparatus comprises: a microphone configured to generate voice signals in response to the voices of the plurality of speakers; a communication circuit configured to transmit and receive data; memory; and a processor, wherein the processor, on the basis of instructions stored in the memory, performs sound source separation of the voice signals on the basis of sound source positions of each of the voices, generates separate voice signals associated with each of the voices according to the sound source separation, determines output modes corresponding to the sound source positions of each of the voices, and uses the communication circuit to output the separate voice signals according to the determined output modes.


