Microphone Array Overlapping Speech Detection via Phase Difference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing overlapping speech detection technologies are primarily designed for single-channel audio and fail to accurately identify overlapping speech in multi-channel audio scenarios using microphone arrays, leading to reduced accuracy and inability to meet product-level detection requirements.
Innovation Solution
An audio signal processing method that utilizes a microphone array to acquire audio signals, generate spatial distribution information of sound sources based on phase difference information, and identify overlapping speech by combining this information with a conversion relationship learned from historical audio signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing overlapping speech detection technology is directly applied to multi-channel audio scenarios using microphone arrays, then the detection process can be implemented, but the accuracy is reduced and product-level detection requirements are not met
Solution Approach 1:
The patent transitions from single-channel audio detection to multi-channel audio detection by incorporating spatial dimension information. It uses microphone arrays to capture audio signals from multiple spatial positions and introduces spatial distribution information as an additional dimension for overlapping speech detection, thereby improving accuracy in multi-channel scenarios
Solution Approach 2:
The patent changes the detection parameters by introducing spatial distribution information alongside temporal conversion relationships. It combines both spatial parameters (from microphone arrays) and temporal parameters (from historical audio signals) to create a more comprehensive detection model that adapts to multi-channel audio scenarios while maintaining high accuracy
2Measurement precision
If single-channel audio detection methods are used, then the system complexity is low, but the detection accuracy in multi-speaker scenarios is insufficient
Solution Approach 1:
The patent makes the microphone array system multi-functional by using it not only for capturing audio signals but also for extracting spatial distribution information. This spatial information is then integrated with temporal conversion relationships to perform both speech detection and spatial analysis, reducing the need for separate systems while improving accuracy
Solution Approach 2:
The patent introduces spatial distribution information as an intermediary that bridges the gap between multi-channel audio input and overlapping speech detection output. This intermediary element processes the spatial characteristics from microphone arrays and combines them with temporal patterns to achieve accurate detection without requiring overly complex system architecture
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The method effectively improves the accuracy of identifying overlapping speech in multi-channel audio scenarios, meeting product-level detection requirements by leveraging spatial distribution information from microphone arrays.
Implementation Method 1
generating spatial distribution information of a current sound source corresponding to the current audio signal based on phase difference information of the current audio signal captured by the at least two microphones
Data Source
AI summary
Audio signal processing methods, systems, terminal devices, conference devices, teaching devices, intelligent vehicle-mounted devices, server device, and computer-readable storage media are provided. The method comprises: obtaining current audio signals acquired by a microphone array, the microphone array comprising at least two microphones; generating, according to phase difference information of the current audio signals acquired by the at least two microphones, current sound source spatial distribution information corresponding to the current audio signals; and according to the current sound source spatial distribution information, in combination with the conversion relationship between single speech and overlapping speech learned on the basis of historical audio signals, identifying whether the current audio signals are overlapping speech. Compared with single-channel audio, the audio signals acquired by the microphone array are used, and the sound source spatial distribution information is included, thus, the techniques of the present disclosure accurately identify whether the current audio signals are overlapping speech, thereby satisfying the detection requirement for a product level.


