Voice Processing Device Using Position-Language Information
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In environments with multiple speakers, existing technologies face challenges in efficiently separating and identifying individual voice signals and determining the language of each speaker, which is time-consuming and resource-intensive, especially when speakers speak simultaneously or in different languages.
Innovation Solution
A voice processing device that uses a microphone to generate voice signals, a processor to determine speaker positions and languages, and a memory to store position-language information, allowing for the separation and translation of voice signals into different languages, thereby generating translated meeting minutes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If voice signals from multiple speakers are received simultaneously, then all voices are captured, but separation and identification of individual speaker voices becomes complex and resource-intensive
Solution Approach 1:
The patent segments the mixed voice signal into individual speaker components by determining sound source positions and separating voices based on spatial information. The processor divides the composite voice signal into multiple independent speaker voice signals using position-language information and sound source separation techniques.
Solution Approach 2:
The patent introduces position-language information as an intermediary element that links sound source positions with language characteristics. This intermediary enables the system to identify and separate individual speaker voices by using spatial position data as a mediating factor in the voice separation process.
2Measurement precision
If language identification is performed using only voice characteristics, then language detection is achieved, but it takes much time and requires many resources
Solution Approach 1:
The patent performs preliminary determination of sound source positions before language identification. By first establishing the spatial position of each speaker and retrieving pre-stored position-language information, the system prepares the language identification process in advance, avoiding time-consuming analysis of voice characteristics alone.
Solution Approach 2:
The patent uses sound source position as an intermediary to facilitate language identification. Instead of directly analyzing voice characteristics to determine language, the system first determines position, then uses position-language information to identify the language, significantly reducing processing time and resource requirements.
3Adaptability or versatility
If translation is performed for multiple speakers in different languages, then translation results are generated, but the process requires significant time and computational resources
Solution Approach 1:
The patent segments the translation process by first separating speaker voices spatially, then identifying languages based on position-language information, and finally translating each speaker's content independently. This segmented approach allows parallel processing of multiple language translations, improving overall efficiency compared to processing all voices together.
Solution Approach 2:
The patent performs preliminary language identification based on sound source positions before the actual translation process. By determining which language each speaker is using in advance through position-language information, the system prepares translation parameters beforehand, enabling more efficient execution of the translation process for multiple speakers.
Data Source
AI summary
A voice processing device for generating translation results for voices of speakers is disclosed. The voice processing device comprises: a microphone for generating voice signals associated with voices of speakers in response to the voices of the speakers; a memory for storing location-language information indicating languages corresponding to sound source locations of the voices of the speakers; and a processor which uses the voice signals and the location-language information so as to generate translation results obtained by translating the languages of the voices of each speaker, and which uses the translation results so as to generate translation conference minutes including the voice contents of each speaker expressed in different languages.


