Distributed Microphone Noise Cancellation for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice-driven systems in noisy environments struggle to distinguish between intended user speech and extraneous background sounds, leading to erroneous interpretations and reduced effectiveness.
Innovation Solution
The system employs both user-worn microphones and environmental microphones to differentiate between user speech and background noise by comparing sound origins, allowing for digital subtraction of background noise and synchronization of asynchronous audio signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple microphones are used to distinguish user speech from background noise, then speech recognition accuracy is improved, but device complexity increases
Solution Approach 1:
The system divides the microphone network into two functional segments: user-worn microphones (first microphones) that capture speech near the user's mouth, and environmental microphones (second microphones) positioned throughout the workspace that capture background noise. This segmentation allows the system to process and compare signals from different locations to distinguish user speech from environmental noise, resolving the contradiction by organizing complexity into functional groups that serve specific purposes.
Solution Approach 2:
The system introduces a server as an intermediary that receives audio signals from multiple microphones, performs signal processing to identify and subtract background noise, and determines whether sounds originate from the user or the environment. This intermediary component centralizes the complex noise-cancellation processing, allowing individual microphones to remain simple while achieving sophisticated noise rejection through coordinated processing.
2Measurement precision
If environmental microphones are deployed throughout the workspace, then background noise identification is improved, but system cost and complexity increase
Solution Approach 1:
The environmental microphones serve multiple functions: they capture background noise for subtraction from user speech signals, they independently detect and identify impulse sounds in the environment, and they provide spatial audio data for determining sound origins. This multi-functionality justifies the deployment of distributed microphones throughout the workspace, as each microphone contributes to multiple aspects of noise cancellation and environmental monitoring.
Solution Approach 2:
The system continuously monitors audio signals from environmental microphones and uses this feedback to dynamically adjust noise cancellation processing. When impulse sounds are detected by environmental microphones, the system responds by subtracting these identified noises from user speech signals in real-time, creating a closed-loop system that adapts to changing environmental acoustic conditions.
3Measurement precision
If digital subtraction of background noise is performed, then speech clarity is improved, but processing complexity increases
Solution Approach 1:
The system performs preliminary identification and isolation of background noise signals before subtracting them from user speech. Environmental microphones continuously capture and analyze environmental sounds, pre-processing the noise data so that when noise cancellation is needed, the system can quickly subtract pre-identified noise patterns from user speech signals, reducing the real-time processing burden.
Solution Approach 2:
The system replaces traditional mechanical or hardware-based noise cancellation approaches with digital signal processing. Instead of using physical acoustic barriers or passive noise isolation, the system uses digital algorithms to identify, analyze, and subtract background noise from audio signals, enabling more flexible and adaptive noise cancellation that can be implemented through software processing.
4Reliability
If impulse sound detection is implemented, then false speech recognition is reduced, but processing time increases
Solution Approach 1:
The system implements periodic monitoring of audio signals for impulse sound detection, analyzing environmental microphone data at regular intervals to identify sudden noise events. This periodic approach allows the system to maintain readiness for impulse detection without continuously processing all audio data at maximum intensity, balancing reliability with processing efficiency.
Solution Approach 2:
The system applies impulse detection processing selectively based on local conditions: when environmental microphones detect characteristics consistent with impulse sounds (sudden, sharp noise events), the system activates enhanced processing for those specific time periods and spatial locations, while using standard processing during normal conditions. This localized application of impulse detection reduces overall processing time while maintaining reliability when needed.
Data Source
AI summary
A device, system, and method whereby a speech-driven system used in an industrial environment distinguishes speech obtained from users of the system from other background sounds. In one aspect, the present system and method provides for a first audio stream from a user microphone collocated with a source of human speech (that is, a user) and a second audio stream from a environmental microphone which is proximate to the source of human speech but more remote than the user microphone. The audio signals from the two microphones are asynchronous. A processor is configured to identify a common, distinctive sound event in the environment, such as an impulse sound or a periodic sound signal. Based on the common sound event, the processor provides for synchronization of the two audio signals. In another aspect, the present system and method provides for a determination of whether or not the sound received at the user microphone is suitable for identification of words in a human voice, based on a comparison of sound elements in the first audio stream and the second audio stream, for example based on a comparison of the sound intensities of the sound elements in the audio streams.


