Speech Signal Processing Using Facial Recognition and Microphone Array
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems face challenges in noisy environments, where voice activity detection (VAD) fails due to interference from multiple sources, leading to reduced accuracy and increased computation load.
Innovation Solution
The method integrates facial recognition with sound source localization and voice activity detection using a microphone array to identify and isolate speech signals in noisy environments, improving the anti-interference capability by determining speech sound segments through computer vision and sound source orientation analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech sound energy detection is used for voice activity detection, then the system can detect speech segments, but it fails in noisy environments with multiple interference sources, reducing accuracy
Solution Approach 1:
The patent segments the sound field into multiple spatial regions using a microphone array to form directional beams. By dividing the monitoring space into different sectors and independently processing signals from each sector, the system can identify and isolate the target speech source from interfering sources located in different directions, thereby improving detection accuracy in noisy environments.
Solution Approach 2:
The patent applies local quality enhancement by optimizing signal processing parameters for specific spatial directions. The beamforming algorithm adjusts the weight coefficients of individual microphones based on the desired direction, creating directional sensitivity that enhances the target speech signal while suppressing noise from other directions, thus improving local signal quality.
2Productivity
If traditional voice activity detection is performed without spatial filtering, then processing is simpler, but computation load increases and response time decreases in noisy environments
Solution Approach 1:
The patent performs preliminary spatial filtering and beamforming operations before voice activity detection. By pre-processing the multi-channel microphone signals to form directional beams and isolate potential speech sources, the system reduces the complexity of subsequent VAD processing and accelerates detection response time, as the search space has already been narrowed down spatially.
3Reliability
If multiple microphones are used for sound source localization, then spatial filtering capability is improved, but device complexity increases
Solution Approach 1:
The patent makes the microphone array system multi-functional by integrating both sound source localization and voice activity detection capabilities into a single unified framework. The same array geometry and signal processing infrastructure serve dual purposes: determining the spatial location of speech sources and simultaneously performing directional speech detection, thereby reducing overall system complexity despite using multiple microphones.
Data Source
AI summary
The disclosed embodiments disclose methods, apparatuses, systems, devices and computer-readable storage media for processing speech signals. The method comprises: acquiring a real-time image by using an image capturing device, performing facial recognition by using the real-time image, and detecting a period during which a target user makes speech sounds based on a facial recognition result; locating a sound source in an audio signal received by a microphone array, and determining the orientation information of a sound source in the audio signal; and based on the period during which the target user in the real-time image makes the speech sounds and the orientation information of the sound source, performing a speech sound start and end point analysis to determine start and end time points of the speech sounds in the audio signal. The method for processing speech signals according to one embodiment can perform voice activity detection to the speech signal in noisy environments containing multiple sources of interference, thereby improving the anti-interference capability of the system.


