Persistent Interference Detection in Smart Home Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Far-field audio processing in smart home devices faces challenges due to reduced intelligibility and interference from noise sources, where the voice amplitude is often drowned out by background noise, leading to incorrect recognition of speech commands.
Innovation Solution
The use of multiple microphones to spatially detect and filter out persistent interference sources by analyzing the inter-microphone frequency-dependent phase profiles, distinguishing them from desired talkers based on their stationary versus moving nature, and enhancing signal processing to improve voice quality and automatic speech recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If far-field audio processing is used to capture speech from distant talkers, then the coverage area is improved, but the speech intelligibility deteriorates due to reduced amplitude and increased reverberation
Solution Approach 1:
The patent combines signals from multiple microphones (at least two microphones) to process far-field speech. By merging the captured audio signals and applying beamforming techniques, the system enhances the speech signal from the desired direction while suppressing reverberation and noise, thereby maintaining speech intelligibility across a wide coverage area
Solution Approach 2:
The patent introduces spatial dimensionality by using multiple microphones arranged in specific geometries (e.g., linear array, triangular array). This spatial arrangement enables the system to distinguish between direct speech paths and reverberant paths based on time-of-arrival differences and phase relationships, improving speech intelligibility in far-field conditions
2Ease of operation
If the microphone is placed far from the talker to enable far-field processing, then the natural interaction is improved, but the speech amplitude decreases making it difficult to discern from noise
Solution Approach 1:
The system merges signals from multiple microphones to achieve coherent integration of the speech signal while incoherent integration of noise and reverberation. This signal combining approach, combined with beamforming, amplifies the desired speech signal relative to background noise, maintaining reliable speech detection even at far distances where natural interaction occurs
Solution Approach 2:
The patent employs rapid processing techniques to extract and enhance the speech signal before it is completely overwhelmed by reverberation and noise. By processing the audio signals in real-time with beamforming and speech enhancement algorithms, the system retrieves the speech content through the noisy reverberant field
3Adaptability or versatility
If interference sources are present in the environment, then the environmental monitoring capability is improved, but the speech recognition accuracy deteriorates due to amplitude competition between speech and interference
Solution Approach 1:
The patent applies beamforming to create directional sensitivity patterns that enhance speech from specific directions while suppressing interference from other directions. By adjusting the beamforming weights locally for different spatial locations, the system can selectively enhance the desired speech signal and suppress interference sources, maintaining high speech recognition accuracy even in environments with multiple active sources
Solution Approach 2:
The system uses spatial filtering and beamforming as intermediary processing steps between raw microphone signals and speech recognition. This intermediary stage separates the desired speech signal from interference by exploiting spatial differences, allowing accurate speech recognition to proceed even when interference sources are present in the environment
4Reliability
If multiple microphones are used to process far-field signals, then the signal-to-noise ratio is improved, but the device complexity increases
Solution Approach 1:
The patent segments the audio processing into distinct functional stages: initial signal capture by individual microphones, beamforming processing to combine signals spatially, and final speech recognition. This segmentation allows each stage to be optimized independently, achieving high signal-to-noise ratio through coordinated processing while managing overall system complexity through modular architecture
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach effectively filters out interference, enhancing the accuracy of speech recognition and reducing false command triggers, thereby improving voice quality and reliability in noisy environments.
Implementation Method 1
an acoustic source is identified as a persistent interference source when the source is determined to be originating from the same spatial location with respect to the microphone array over several time periods based on an inter-microphone frequency-dependent phase profile
Data Source
AI summary
A multi-microphone algorithm for detecting and differentiating interference sources from desired talker speech in advanced audio processing for smart home applications is described. The approach is based on characterizing a persistent interference source when sounds repeated occur from a fixed spatial location relative to the device, which is also fixed. Some examples of such interference sources include TV, music system, air-conditioner, washing machine, and dishwasher. Real human talkers, in contrast, are not expected to remain stationary and speak continuously from the same position for a long time. The persistency of an acoustic source is established based on identifying historically-recurring inter-microphone frequency-dependent phase profiles in multiple time periods of the audio data. The detection algorithm can be used with a beamforming processor to suppress the interference and for achieving voice quality and automatic speech recognition rate improvements in smart home applications.


