Directional Audio Beamformer for False Trigger Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Devices that interact with users through speech struggle to isolate user speech in environments with interfering sounds from media sources like televisions and radios, leading to false detection of trigger expressions.

Innovation Solution

The audio device employs an audio beamformer to produce directional audio signals, a speech activity detector to identify human speech, and a sound source detector to differentiate between human and electronic sources, applying different detection standards to reduce false positives from electronic sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the speech activity detector detects all speech-like sounds in the environment, then the sensitivity of expression detection is improved, but false detections from electronic sources increase

Engineering Contradiction:
Improveexpression detection sensitivityVSAvoidfalse detection rate
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the audio signal processing into distinct functional components: a speech activity detector that identifies potential speech, a sound source detector that analyzes signal characteristics to determine origin, and an expression detector that performs final analysis. This segmentation allows each component to specialize, with the sound source detector filtering out electronic sources before the expression detector processes the signal, thus maintaining sensitivity while reducing false detections.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The sound source detector acts as an intermediary component between the speech activity detector and the expression detector. It receives directional audio signals, analyzes their characteristics to determine whether they originate from human or electronic sources, and selectively provides signals to the expression detector. This intermediary filtering mechanism resolves the contradiction by allowing high sensitivity in detecting human speech while blocking false detections from electronic sources.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the system applies strict detection standards to filter electronic sources, then false detections are reduced, but the detection of legitimate human speech may be missed

Engineering Contradiction:
Improvefalse detection rateVSAvoidexpression detection sensitivity
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent applies different detection standards locally based on the identified sound source type. The sound source detector analyzes signal characteristics and selectively applies filtering criteria: stringent filtering for signals identified as electronic sources, and permissive detection for signals identified as human speech. This local differentiation resolves the contradiction by maintaining high reliability for electronic source rejection while preserving sensitivity for legitimate human speech detection.

Inventive Principle:
Principle #3Local quality

3Reliability

If the system analyzes signal characteristics to identify electronic sources, then false detections are reduced, but the device complexity increases

Engineering Contradiction:
Improvefalse detection rateVSAvoidsignal processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The sound source detector performs preliminary analysis of signal characteristics before the expression detector processes the signal. By pre-identifying electronic sources through characteristic analysis and filtering them out in advance, the system reduces the processing burden on subsequent stages. This preliminary action resolves the contradiction by establishing reliability through early filtering while managing complexity through staged processing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9734845B1Mitigating effects of electronic audio sources in expression detection
Publication Date: 2017.08.15 AMAZON TECH INC
  • US9734845B1 patent drawing
  • US9734845B1 patent drawing
  • US9734845B1 patent drawing

AI summary

In a speech-based system, a wake word or other trigger expression is used to preface user speech that is intended as a command. The system receives multiple directional audio signals, each of which emphasizes sound from a different direction. The signals are monitored and analyzed to detect the directions of interfering audio sources such as televisions or other types of electronic audio players. One of the directional signals having the strongest presence of speech is selected to be monitored for the trigger expression. If the directional signal corresponds to the direction of an interfering audio source, a more strict standard is used to detect the trigger expression. In addition, the directional audio signal having the second strongest presence of speech may also be monitored to detect the trigger expression.