3-Tuple Coprime Microphone Array for Multi-Talker Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing microphone arrays face challenges in effectively separating multiple talkers due to broad beam widths and increased complexity with more microphones, particularly at higher frequencies, where grating lobes introduce aliasing and make it difficult to accurately locate sound sources.
Innovation Solution
A 3-tuple coprime microphone array is used, comprising three uniform linear subarrays with coprime numbers of elements, which collectively beamform signals through point-by-point multiplication and cube-root operations to achieve directional filtering and separate speech signals from noise, allowing for beam steering and efficient sound source separation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If more microphones are used in the array, then the beam width narrows and sound source localization improves, but the device complexity increases and computational requirements increase
Solution Approach 1:
The microphone array is divided into three separate uniform linear subarrays with coprime numbers of elements. Each subarray independently processes signals and then combines through point-by-point multiplication, achieving the localization precision of a large array while using fewer total microphones and reducing system complexity
Solution Approach 2:
The patent transitions from a two-dimensional uniform array to a three-dimensional coprime subarray configuration. By adding the temporal dimension through point-by-point multiplication of subarray outputs, the system achieves enhanced directional resolution without proportionally increasing the number of microphones
2Measurement precision
If the operating frequency is increased, then the beam width narrows and directional resolution improves, but grating lobes appear causing aliasing and making sound source location difficult
Solution Approach 1:
The patent changes the spatial sampling parameters by using coprime subarray configurations with different element spacings. This parameter variation ensures that grating lobes from different subarrays occur at different angles, allowing the main lobe to be distinguished from aliasing artifacts even at high frequencies
Solution Approach 2:
The point-by-point multiplication operation acts as an intermediary that combines subarray outputs. This operation suppresses grating lobe artifacts while preserving the main lobe signal, effectively filtering out aliasing effects and enabling accurate sound source localization at high frequencies
3Measurement precision
If a larger microphone array is used to reduce beam width, then multi-talker separation improves, but the computational complexity increases
Solution Approach 1:
The signal processing is segmented into independent subarray processing stages followed by a simple combination operation. Each subarray processes signals independently and then combines through point-by-point multiplication, reducing computational complexity compared to processing a large unified array while maintaining multi-talker separation capability
Data Source
AI summary
A method of multi-talker separation using a 3-tuple coprime microphone array, including generating, by a subarray signal processing module, a respective subarray data set for each microphone subarray of the 3-tuple coprime microphone array based, at least in part, on an input acoustic signal comprising at least one speech signal. The input acoustic signal is captured by the 3-tuple coprime microphone array. The 3-tuple coprime microphone array includes three microphone subarrays. The method includes determining, by the subarray signal processing module, a point by point product of the three subarray data sets; and determining, by the subarray signal processing module, a cube root of the point by point product to yield an acoustic signal output data. The acoustic signal output data has an output amplitude and an output phase corresponding to an input amplitude and an input phase of a selected speech signal of the at least one speech signal.


