Vehicle Voice Processing Using Spatial Separation for Fast Translation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice processing systems struggle to efficiently separate voice signals from multiple speakers in a vehicle and translate them into different languages without requiring extensive time and resources to identify the source languages.

Innovation Solution

A voice processing device equipped with a voice processing circuit for separating voice signals based on voice source positions, a memory for storing language information, and a communication circuit for outputting translations, which determines source and target languages from voice source positions to provide translations efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice source separation is performed based on voice source positions, then separated voice signals for individual speakers are generated, but the system complexity increases due to the need for position determination and separation processing

Engineering Contradiction:
Improvevoice separation accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the mixed voice signal into individual speaker signals by separating voices based on their spatial positions. The voice processing circuit divides the composite voice signal received from multiple speakers into separate voice signals corresponding to each speaker's location, thereby achieving accurate voice separation through spatial segmentation.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If source language identification is performed for each speaker, then accurate translation can be provided, but the time and resources required increase significantly

Engineering Contradiction:
Improvelanguage identification accuracyVSAvoidtranslation processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by determining the voice source positions of each speaker before the translation process. The system determines the positions of multiple voice sources and uses this spatial information to directly associate each separated voice signal with its corresponding speaker, eliminating the need for time-consuming language identification and enabling immediate translation processing.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If multiple speakers pronounce in different languages, then translation capability is required, but the resource requirements increase to grasp the languages of each speaker

Engineering Contradiction:
Improvemulti-language translation capabilityVSAvoidprocessing resources
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent introduces voice source position as an intermediary element that mediates between the mixed voice signal and the translation process. By determining the positions of multiple voice sources and using these positions to separate and identify speaker signals, the system enables multi-language translation without requiring extensive resources for language identification, as the spatial information serves as a efficient mediator.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12347450B2Voice processing device and operating method therefor
Publication Date: 2025.07.01 AMOSENSE CO LTD
  • US12347450B2 patent drawing
  • US12347450B2 patent drawing
  • US12347450B2 patent drawing

AI summary

A voice processing device is disclosed. The voice processing device comprises: a voice processing circuit configured to generate an isolated voice signal associated with respective voices spoken at a plurality of sound source locations in a vehicle by isolating, and output an interpretation result for the respective voices on the basis of the isolated voice signal; a memory configured to store source language information indicating a source language and target language information indicating a target language in order to interpret the voice associated with the isolation voice signal; and a communication circuit configured to output the interpretation result, wherein the voice processing circuit generates the interpretation result in which the language of the voice corresponding to the isolated voice signal is interpreted from the source language into the target language.