Spatial Audio Database Noise Discrimination for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound recognition devices face significant performance degradation in recognizing speech commands when the speaker is not close to the device or when background noise is present, leading to user frustration and potential safety issues.
Innovation Solution
The implementation of a spatial audio database-based noise discrimination method, which uses an array of microphones and a digital signal processor to segregate sounds into spatial signals, compare them to a spatial audio database, and apply a spatial-temporal filter to distinguish speech commands from background noise, thereby improving recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional sound recognition methods are used, then the device can recognize speech commands in ideal conditions, but performance degrades significantly when background noise is present or speaker is far away
Solution Approach 1:
The patent segments the audio signal into multiple spatial components using an array of microphones and signal processing techniques. Each microphone captures sound from different spatial positions, and the system separates these spatial signals to identify and isolate speech commands from background noise sources based on their spatial characteristics and temporal patterns.
Solution Approach 2:
The patent introduces a spatial-temporal filter as an intermediary processing layer between the microphones and the speech recognition system. This filter acts as a mediator that selectively enhances speech signals while suppressing background noise based on spatial and temporal characteristics, improving the signal-to-noise ratio before recognition.
2Ease of operation
If the speaker is positioned far from the device, then user convenience is improved, but speech command recognition accuracy deteriorates
Solution Approach 1:
The patent combines signals from multiple microphones in a spatial array to create a composite audio signal with enhanced spatial information. By merging the captured signals and processing them together through spatial-temporal filtering, the system maintains recognition accuracy even when the speaker is at a distance, as the combined signal provides better spatial discrimination capability.
3Measurement precision
If spatial-temporal filtering is applied to filter out background noise, then speech recognition accuracy is improved, but device complexity increases
Solution Approach 1:
The patent performs preliminary spatial signal separation and noise discrimination before the main speech recognition process. By pre-processing the audio signals to isolate speech components from background noise using spatial-temporal characteristics, the system simplifies the subsequent recognition task and improves overall accuracy without requiring overly complex real-time processing during speech recognition.
Data Source
AI summary
Methods, systems, and computer-readable and executable instructions for spatial audio database based noise discrimination are described herein. For example, one or more embodiments include comparing a sound received from a plurality of microphones to a spatial audio database, discriminating a speech command and a background noise from the received sound based on the comparison to the spatial audio database, and determining an instruction based on the discriminated speech command.


