Voice Extraction via Microphone Array Feature Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech extraction methods in intelligent hardware, such as smart speakers and in-vehicle devices, face challenges in accurately isolating speech signals from noise and reverberation, leading to reduced accuracy and reliability in speech recognition systems.
Innovation Solution
A method and apparatus that utilize a microphone array to obtain and process data, employing signal processing techniques like beamforming and blind source separation to extract normalized features characterizing speech presence, and fuse these features with speech features to enhance speech data extraction in specific directions, improving accuracy and reducing environmental noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional speech extraction methods are used in intelligent hardware, then the system structure remains simple, but the speech extraction accuracy deteriorates due to noise and reverberation
Solution Approach 1:
The patent segments the speech extraction process into multiple independent modules: beamforming module for spatial filtering, blind source separation module for source decomposition, and feature fusion module for combining multiple features. This segmentation allows each module to specialize in specific noise reduction tasks while maintaining overall system manageability and improving extraction accuracy through coordinated operation of separate functional units.
Solution Approach 2:
The patent employs a composite processing approach by fusing normalized features from beamforming with speech features from blind source separation. This fusion of multiple feature types (spatial features, spectral features, temporal features) creates a composite feature representation that leverages the strengths of different processing methods to achieve superior speech extraction accuracy compared to single-method approaches.
2Reliability
If advanced signal processing techniques are applied to reduce noise and reverberation, then speech extraction accuracy improves, but the processing complexity increases
Solution Approach 1:
The patent applies beamforming as a preliminary processing step before blind source separation. The beamforming module pre-processes the microphone array data to create spatially-filtered signals that emphasize speech directions and suppress noise, providing cleaner input to subsequent blind source separation and feature extraction stages. This preliminary action reduces the burden on later processing stages and improves overall reliability.
Solution Approach 2:
The patent implements feature fusion that combines normalized features (from beamforming) with speech features (from blind source separation) in a feedback loop. The fused features are used to iteratively refine speech extraction, where the output of one processing stage feeds back to enhance the input of the next stage, allowing the system to adaptively improve speech recognition reliability through multiple processing passes.
3Measurement precision
If microphone array data is processed using multiple signal processing techniques, then the speech presence characterization improves, but the processing time increases
Solution Approach 1:
The patent applies partial processing by selectively applying different signal processing techniques to different frequency bands or time segments of the audio signal. Instead of processing the entire signal through all processing stages uniformly, the system applies beamforming and blind source separation only to portions of the signal where speech is detected or where noise suppression is most needed, reducing overall processing time while maintaining speech presence characterization accuracy in critical regions.
Data Source
AI summary
A voice extraction method and apparatus (500), and an electronic device. The method comprises: acquiring microphone array data (303) (201, 401); performing signal processing on the microphone array data (303) to obtain a normalized feature (304) (202, 402), wherein the normalized feature (304) is used for representing the probability of a voice being present in a predetermined direction; on the basis of the microphone array data (303), determining a voice feature (306) of a voice in a target direction (203); and fusing the normalized feature (304) with the voice feature (306) of the voice in the target direction, and extracting voice data (309) in the target direction according to the voice feature (307) after same is subjected to fusion (204). Environmental noise is reduced, and the accuracy of extracted voice data is improved.


