Microphone Array Spatial Segregation for Distant Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Previous sound recognition devices struggle to accurately recognize speech commands when the speaker is distant or in the presence of background noise, such as from other speakers, televisions, appliances, or animals.
Innovation Solution
The system employs an array of microphones and a digital signal processor to spatially segregate sound into multiple signals, using beam former algorithms to isolate speech commands from background noise, allowing an automatic speech recognition engine to process each signal separately and improve recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a single microphone is used for sound recognition, then the device structure is simple, but the recognition accuracy degrades significantly when the speaker is distant or background noise is present
Solution Approach 1:
The patent divides the sound capture function into multiple microphones arranged in an array, with each microphone capturing sound from different spatial positions. The digital signal processor then segments the captured sound into multiple directional signals, allowing the system to identify and process speech from specific directions while filtering out noise from other directions.
Solution Approach 2:
The patent transitions from a single-point sound capture (one microphone) to a spatial distribution of sound capture (microphone array). By adding the spatial dimension with multiple microphones positioned at different locations, the system can distinguish speech from noise based on directional information and spatial characteristics.
2Reliability
If beam former algorithms are used to spatially segregate sound signals, then speech recognition in noisy environments improves, but computational complexity increases
Solution Approach 1:
The patent applies beam former algorithms to pre-process the captured sound signals before they reach the speech recognition engine. By performing spatial segmentation and noise filtering in advance, the system prepares cleaner, directionally-separated signals for the recognition engine, improving reliability while managing computational complexity through staged processing.
3Measurement precision
If multiple microphones and signal processing algorithms are implemented, then speech recognition performance in distant and noisy conditions improves, but the device cost and complexity increase
Solution Approach 1:
The system segments both the physical microphone array and the signal processing tasks. Multiple microphones are distributed spatially, and the digital signal processor divides the audio spectrum and spatial information into separate channels, processing each segment independently to achieve high distant speech recognition accuracy while managing overall system complexity.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables effective recognition of speech commands from a distance and in noisy environments, enhancing the performance of sound recognition devices by isolating speech from background noise, thereby improving their operational reliability.
Implementation Method 1
a digital signal processor configured to segregate (e.g., spatially segregate) the captured sound into a plurality of signals, wherein each respective signal corresponds to a different portion of the area
Data Source
AI summary
Speech recognition methods, devices, and systems are described herein. One system includes a number of microphones configured to capture sound in an area, a digital signal processor configured to segregate the captured sound into a plurality of signals, wherein each respective signal corresponds to a different portion of the area, and an automatic speech recognition engine configured to separately process each of the plurality of signals to recognize a speech command in the captured sound.

