In-Vehicle Voice Processing for Conversation-Aware Call Areas
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing in-vehicle systems struggle to effectively manage voice call areas based on the conversation mode, failing to accurately identify and switch call areas according to the utterers present in different seats within a vehicle.
Innovation Solution
A voice processing device that utilizes multiple microphones and speakers to recognize voices in specific areas, associate area information with utterer information, and selectively switch call areas based on conversation mode settings, employing beam forming, echo and cross-talk cancellation, and speaking person recognition to ensure clear voice communication.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple microphones and speakers are used to recognize voices in different areas, then voice recognition accuracy is improved, but device complexity increases
Solution Approach 1:
The vehicle interior is divided into multiple call areas (front seat area, rear seat area) with dedicated microphones and speakers for each area. This segmentation enables independent voice recognition and call area switching for different seating positions, improving voice recognition accuracy while managing complexity through modular area-based processing
Solution Approach 2:
The system adds a spatial dimension to voice processing by implementing beam forming that processes voice signals from multiple microphones to determine the direction and area of utterers. This dimensional approach enables accurate call area identification without requiring complex individual speaker recognition for each position
2Adaptability or versatility
If call area switching based on conversation mode is implemented, then adaptability is improved, but device complexity increases
Solution Approach 1:
The system dynamically switches call areas based on detected conversation modes (one-to-one, one-to-many). The control unit automatically adjusts which microphones and speakers are active according to the conversation scenario, providing adaptability without requiring complex manual configuration
Solution Approach 2:
The same microphone and speaker array serves multiple functions: voice recognition, call area determination, and adaptive switching for different conversation modes. This multi-functionality reduces the need for separate dedicated systems for each mode, managing complexity while maintaining versatility
3Reliability
If beam forming and echo cancellation are used to maintain voice quality, then voice communication quality is improved, but processing time increases
Solution Approach 1:
The beam forming and echo cancellation parameters are pre-configured for each call area and conversation mode. When switching between areas or modes, the system activates pre-prepared signal processing parameters, reducing the time required for real-time calculation while maintaining voice communication quality
Data Source
AI summary
A voice processing device according to the present disclosure includes a memory in which a program is stored, and a processor coupled to the memory and configured to perform processing by executing the program. The processing includes processing voice signals input from a plurality of microphones; recognizing voices of utterers each present in corresponding one of areas, based on the voice signals; associating area information of each of the areas with utterer information of corresponding one of the utterers whose voice is recognized in corresponding one of the areas; and selectively switching a call area to a call area corresponding to a setting of a conversation mode, based on the utterer information. In the selectively switching, the call area is selected by selection of a voice signal from among the voice signals and selection of a speaker that outputs the voice signal from among a plurality of speakers.


