Speech Signal Processing Using Facial Recognition and Microphone Array

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems face challenges in noisy environments, where voice activity detection (VAD) fails due to interference from multiple sources, leading to reduced accuracy and increased computation load.

Innovation Solution

The method integrates facial recognition with sound source localization and voice activity detection using a microphone array to identify and isolate speech signals in noisy environments, improving the anti-interference capability by determining speech sound segments through computer vision and sound source orientation analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech sound energy detection is used for voice activity detection, then the system can detect speech segments, but it fails in noisy environments with multiple interference sources, reducing accuracy

Engineering Contradiction:
Improvespeech detection accuracyVSAvoidnoise interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent segments the sound field into multiple spatial regions using a microphone array to form directional beams. By dividing the monitoring space into different sectors and independently processing signals from each sector, the system can identify and isolate the target speech source from interfering sources located in different directions, thereby improving detection accuracy in noisy environments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality enhancement by optimizing signal processing parameters for specific spatial directions. The beamforming algorithm adjusts the weight coefficients of individual microphones based on the desired direction, creating directional sensitivity that enhances the target speech signal while suppressing noise from other directions, thus improving local signal quality.

Inventive Principle:
Principle #3Local quality

2Productivity

If traditional voice activity detection is performed without spatial filtering, then processing is simpler, but computation load increases and response time decreases in noisy environments

Engineering Contradiction:
Improveprocessing efficiencyVSAvoiddetection response time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary spatial filtering and beamforming operations before voice activity detection. By pre-processing the multi-channel microphone signals to form directional beams and isolate potential speech sources, the system reduces the complexity of subsequent VAD processing and accelerates detection response time, as the search space has already been narrowed down spatially.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If multiple microphones are used for sound source localization, then spatial filtering capability is improved, but device complexity increases

Engineering Contradiction:
Improveanti-interference capabilityVSAvoidmicrophone array complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent makes the microphone array system multi-functional by integrating both sound source localization and voice activity detection capabilities into a single unified framework. The same array geometry and signal processing infrastructure serve dual purposes: determining the spatial location of speech sources and simultaneously performing directional speech detection, thereby reducing overall system complexity despite using multiple microphones.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11398235B2Methods, apparatuses, systems, devices, and computer-readable storage media for processing speech signals based on horizontal and pitch angles and distance of a sound source relative to a microphone array
Publication Date: 2022.07.26 ALIBABA GROUP HOLDING LTD
  • US11398235B2 patent drawing
  • US11398235B2 patent drawing
  • US11398235B2 patent drawing

AI summary

The disclosed embodiments disclose methods, apparatuses, systems, devices and computer-readable storage media for processing speech signals. The method comprises: acquiring a real-time image by using an image capturing device, performing facial recognition by using the real-time image, and detecting a period during which a target user makes speech sounds based on a facial recognition result; locating a sound source in an audio signal received by a microphone array, and determining the orientation information of a sound source in the audio signal; and based on the period during which the target user in the real-time image makes the speech sounds and the orientation information of the sound source, performing a speech sound start and end point analysis to determine start and end time points of the speech sounds in the audio signal. The method for processing speech signals according to one embodiment can perform voice activity detection to the speech signal in noisy environments containing multiple sources of interference, thereby improving the anti-interference capability of the system.