Voice Processing Device Speaker Position Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In environments with multiple speakers, existing technologies face challenges in accurately separating and translating voice signals in real-time, especially when speakers speak simultaneously or in different languages, requiring significant time and resources to identify languages and separate voices.
Innovation Solution
A voice processing device and system that uses input voice signals to determine speaker positions, separate voice signals by speaker, and translate languages in real-time, utilizing a processor, memory, and communication circuits to generate and transmit output voice data to translation environments for codeless earphones.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If voice signals from multiple speakers are received simultaneously, then the microphone captures all voices, but it becomes difficult to separate and identify individual speaker voices and their languages
Solution Approach 1:
The patent segments the mixed voice signals into individual speaker components by utilizing spatial information from multiple microphones. Each speaker's voice is separated into distinct signal channels based on their positional characteristics, enabling individual identification and language detection without manual intervention.
Solution Approach 2:
The patent introduces an intermediary processing system that includes a processor and memory unit. This intermediary component receives the mixed voice signals, performs automated speaker separation and language identification through algorithmic analysis, and outputs structured information about each speaker's language and voice characteristics.
2Measurement precision
If manual language identification is performed, then source languages can be determined, but it requires significant time and computational resources
Solution Approach 1:
The patent implements self-service language identification where the system automatically detects and identifies the source languages of speakers without requiring manual intervention. The processor analyzes voice signal characteristics and autonomously determines language information, storing results in the memory unit for subsequent processing.
Solution Approach 2:
The patent performs preliminary language identification and speaker separation actions before the main translation processing. By pre-identifying languages and separating voices in advance, the system prepares structured output data that can be directly used in translation operations, eliminating the need for time-consuming manual language detection during translation.
3Quantity of substance
If all voice signals are processed together, then complete voice data is available, but translation accuracy decreases due to inability to distinguish individual speakers
Solution Approach 1:
The patent segments the complete voice signal into individual speaker components before translation. By dividing the mixed audio into separate channels corresponding to each speaker, the system maintains complete voice data while enabling precise translation of each speaker's language individually, significantly improving translation accuracy.
Data Source
AI summary
A voice processing device is disclosed. The voice processing device comprises: a voice data receiving circuit receives input voice data associated with voices of speakers; a memory stores starting language data; a voice data output circuit outputs output voice data associated with the voices of the speakers; and a processor generates a control command for outputting the output voice data, wherein the processor uses the input voice data to generate first speaker position data indicating a position of a first speaker of the speakers and first output voice data associated with a voice of the first speaker, reads first source language data corresponding to the first speaker position data with reference to the memory, and transmits, to the voice data output circuit, a control command for outputting the first output voice data to a translation environment for translating a first source language indicated by the first starting language data.


