Vehicle Voice Processing Using Spatial Separation for Fast Translation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice processing systems struggle to efficiently separate voice signals from multiple speakers in a vehicle and translate them into different languages without requiring extensive time and resources to identify the source languages.
Innovation Solution
A voice processing device equipped with a voice processing circuit for separating voice signals based on voice source positions, a memory for storing language information, and a communication circuit for outputting translations, which determines source and target languages from voice source positions to provide translations efficiently.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If voice source separation is performed based on voice source positions, then separated voice signals for individual speakers are generated, but the system complexity increases due to the need for position determination and separation processing
Solution Approach 1:
The patent segments the mixed voice signal into individual speaker signals by separating voices based on their spatial positions. The voice processing circuit divides the composite voice signal received from multiple speakers into separate voice signals corresponding to each speaker's location, thereby achieving accurate voice separation through spatial segmentation.
2Measurement precision
If source language identification is performed for each speaker, then accurate translation can be provided, but the time and resources required increase significantly
Solution Approach 1:
The patent performs preliminary action by determining the voice source positions of each speaker before the translation process. The system determines the positions of multiple voice sources and uses this spatial information to directly associate each separated voice signal with its corresponding speaker, eliminating the need for time-consuming language identification and enabling immediate translation processing.
3Adaptability or versatility
If multiple speakers pronounce in different languages, then translation capability is required, but the resource requirements increase to grasp the languages of each speaker
Solution Approach 1:
The patent introduces voice source position as an intermediary element that mediates between the mixed voice signal and the translation process. By determining the positions of multiple voice sources and using these positions to separate and identify speaker signals, the system enables multi-language translation without requiring extensive resources for language identification, as the spatial information serves as a efficient mediator.
Data Source
AI summary
A voice processing device is disclosed. The voice processing device comprises: a voice processing circuit configured to generate an isolated voice signal associated with respective voices spoken at a plurality of sound source locations in a vehicle by isolating, and output an interpretation result for the respective voices on the basis of the isolated voice signal; a memory configured to store source language information indicating a source language and target language information indicating a target language in order to interpret the voice associated with the isolation voice signal; and a communication circuit configured to output the interpretation result, wherein the voice processing circuit generates the interpretation result in which the language of the voice corresponding to the isolated voice signal is interpreted from the source language into the target language.


