Directional Audio Beamformer for False Trigger Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Devices that interact with users through speech struggle to isolate user speech in environments with interfering sounds from media sources like televisions and radios, leading to false detection of trigger expressions.
Innovation Solution
The audio device employs an audio beamformer to produce directional audio signals, a speech activity detector to identify human speech, and a sound source detector to differentiate between human and electronic sources, applying different detection standards to reduce false positives from electronic sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the speech activity detector detects all speech-like sounds in the environment, then the sensitivity of expression detection is improved, but false detections from electronic sources increase
Solution Approach 1:
The patent segments the audio signal processing into distinct functional components: a speech activity detector that identifies potential speech, a sound source detector that analyzes signal characteristics to determine origin, and an expression detector that performs final analysis. This segmentation allows each component to specialize, with the sound source detector filtering out electronic sources before the expression detector processes the signal, thus maintaining sensitivity while reducing false detections.
Solution Approach 2:
The sound source detector acts as an intermediary component between the speech activity detector and the expression detector. It receives directional audio signals, analyzes their characteristics to determine whether they originate from human or electronic sources, and selectively provides signals to the expression detector. This intermediary filtering mechanism resolves the contradiction by allowing high sensitivity in detecting human speech while blocking false detections from electronic sources.
2Reliability
If the system applies strict detection standards to filter electronic sources, then false detections are reduced, but the detection of legitimate human speech may be missed
Solution Approach 1:
The patent applies different detection standards locally based on the identified sound source type. The sound source detector analyzes signal characteristics and selectively applies filtering criteria: stringent filtering for signals identified as electronic sources, and permissive detection for signals identified as human speech. This local differentiation resolves the contradiction by maintaining high reliability for electronic source rejection while preserving sensitivity for legitimate human speech detection.
3Reliability
If the system analyzes signal characteristics to identify electronic sources, then false detections are reduced, but the device complexity increases
Solution Approach 1:
The sound source detector performs preliminary analysis of signal characteristics before the expression detector processes the signal. By pre-identifying electronic sources through characteristic analysis and filtering them out in advance, the system reduces the processing burden on subsequent stages. This preliminary action resolves the contradiction by establishing reliability through early filtering while managing complexity through staged processing.
Data Source
AI summary
In a speech-based system, a wake word or other trigger expression is used to preface user speech that is intended as a command. The system receives multiple directional audio signals, each of which emphasizes sound from a different direction. The signals are monitored and analyzed to detect the directions of interfering audio sources such as televisions or other types of electronic audio players. One of the directional signals having the strongest presence of speech is selected to be monitored for the trigger expression. If the directional signal corresponds to the direction of an interfering audio source, a more strict standard is used to detect the trigger expression. In addition, the directional audio signal having the second strongest presence of speech may also be monitored to detect the trigger expression.


