In-Vehicle Voice Separation for Single-Receiver Cabin Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The high cost of configuring multiple voice receivers in vehicles for voice recognition capabilities, which increases the overall cost and data processing complexity.
Innovation Solution
A method that separates initial voice data received from multiple regions within a vehicle into sub-data and corresponding description information, allowing a single voice receiver to determine the vehicle's working mode based on the sub-data, thereby reducing the number of required voice receivers and improving processing performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple voice receivers are configured in different regions of the vehicle, then voice recognition coverage is improved, but device cost and data processing complexity increase
Solution Approach 1:
A single voice receiver is designed to perform multiple functions by receiving and processing voice data from multiple vehicle regions. The receiver uses beamforming technology to identify voice sources from different directions (driver, front passenger, rear passengers) and determines working modes based on the number and location of voice sources, eliminating the need for separate receivers in each region while maintaining comprehensive voice recognition coverage.
2Adaptability or versatility
If multiple voice receivers are configured in different regions of the vehicle, then voice recognition coverage is improved, but device cost increases
Solution Approach 1:
A single voice receiver is designed to perform multiple functions by receiving and processing voice data from multiple vehicle regions. The receiver uses beamforming technology to identify voice sources from different directions (driver, front passenger, rear passengers) and determines working modes based on the number and location of voice sources, eliminating the need for separate receivers in each region while maintaining comprehensive voice recognition coverage.
Solution Approach 2:
Multiple voice reception functions that would traditionally require separate receivers are merged into a single receiver. The receiver combines signal processing capabilities to handle voices from multiple regions simultaneously, using beamforming to spatially separate and identify different voice sources, thereby reducing the total number of receivers from multiple to one while maintaining full coverage.
3Device complexity
If a single voice receiver is used to process voice data from multiple regions, then cost and processing load are reduced, but voice recognition accuracy may deteriorate
Solution Approach 1:
The patent replaces traditional mechanical signal processing methods with beamforming technology, which uses signal processing algorithms to spatially filter and identify voice sources. The beamforming process calculates direction of arrival (DOA) of voice signals and determines the number of voice sources through spectral analysis, enabling accurate identification of voice regions without requiring multiple physical receivers. This substitution maintains high recognition accuracy while reducing system complexity.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method and an apparatus of processing a voice for a vehicle, a device, a medium and a product are provided, which relate to a field of voice recognition technology. The method of processing a voice for a vehicle includes: separating an initial voice data in response to receiving the initial voice data from a plurality of regions inside the vehicle, so as to obtain a plurality of voice sub-data and a description information for each voice sub-data of the plurality of voice sub-data, the plurality of voice sub-data correspond to the plurality of regions respectively, and the description information for each voice sub-data indicates the region corresponding to the each voice sub-data in the plurality of regions; and determining a voice working mode of the vehicle based on the plurality of voice sub-data.