Target Voice Detection Using Microphone Array Beamforming
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing target voice detection methods face limitations in application scenarios and accuracy, particularly in low signal-to-noise environments, and fail to fully utilize multi-channel information effectively.
Innovation Solution
A target voice detection method and apparatus that utilize a microphone array to perform beamforming, extract comprehensive detection features, and employ a pre-constructed model for accurate detection, combining intensity difference-based and model-based results to enhance detection accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If intensity difference-based target voice detection method is used, then the detection process is simple, but the detection accuracy is poor in low signal-to-noise environments and limited application scenarios
Solution Approach 1:
The patent segments the detection process into multiple independent modules: beamforming module for spatial filtering, feature extraction module for extracting multiple types of features (spectral, spatial, temporal), and a detection model that processes these features. This segmentation allows each module to specialize in one aspect, improving overall detection accuracy while maintaining manageable system complexity.
Solution Approach 2:
The patent introduces spatial dimension by using microphone arrays and beamforming techniques to extract spatial features in addition to traditional spectral features. This adds a new dimension (spatial dimension) to the detection process, enabling the system to distinguish target voice from noise based on spatial information, thereby improving detection accuracy in low signal-to-noise environments.
2Measurement precision
If machine learning-based target voice detection method is used, then detection accuracy can be improved, but multi-channel spatial information is not fully utilized and the system performance degrades with human acoustic interference
Solution Approach 1:
The patent segments the feature extraction process into distinct spectral feature extraction and spatial feature extraction modules. The spatial feature extraction module specifically processes multi-channel information from microphone arrays using beamforming, while the spectral feature module handles traditional audio features. This segmentation allows the system to independently optimize each feature type and combine them effectively, improving robustness to interference.
Solution Approach 2:
The patent explicitly adds spatial dimension to the feature space by extracting spatial features from multi-channel microphone inputs through beamforming. This spatial dimension provides additional discriminative information that helps distinguish target voice from human acoustic interference in other directions, making the detection system more adaptable and robust to various interference scenarios.
3Device complexity
If single-channel signal processing is used, then the system complexity is low, but information utilization is insufficient and detection effect is poor
Solution Approach 1:
The patent segments the multi-channel signal processing into independent feature extraction streams for each channel, then combines them through spatial feature processing. This segmentation approach allows the system to process each channel's information separately (maintaining clarity and manageability) while still utilizing all available information from multiple channels, thus reducing information loss without excessive complexity.
Solution Approach 2:
The patent transitions from single-channel processing to multi-channel processing by adding the spatial dimension. The beamforming module processes signals from multiple microphone channels simultaneously, extracting spatial characteristics that are completely lost in single-channel processing. This dimensional expansion fully utilizes multi-channel information while the modular architecture keeps system complexity manageable.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A target voice detection method and apparatus. The method comprises: receiving a sound signal collected on the basis of a microphone array (101); performing beam forming treatment on the sound signal to obtain wave beams in different directions (102); extracting detection features on the basis of the sound signal and the wave beams in different directions frame by frame (103); inputting the extracted detection features of the current frame to a pre-built target voice detection model to obtain a model output result (104); and obtaining a detection result of the target voice corresponding to the current frame according to the model output result (105). Therefore, accuracy of detection results can be improved.