Wake Word Detection Using Contextual Audio Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional wake word detection systems in computing devices often experience high false detection and rejection rates, especially in noisy environments or during ongoing conversations, due to their reliance on acoustic features alone, which can negatively impact user experience.

Innovation Solution

The implementation of a wake word detection system that utilizes a combination of acoustic, environmental, lexical, linguistic, and contextual information, along with probabilistic logic and multiple detectors in series, to make holistic determinations about the presence of a wake word, thereby reducing false detections and rejections.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional wake word detection systems rely on acoustic features alone, then the device complexity is low, but the detection accuracy deteriorates with high false detection and rejection rates in noisy environments

Engineering Contradiction:
Improvewake word detection accuracyVSAvoiddetection system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple detection approaches (acoustic feature detection, keyword spotting, and speech activity detection) into a unified wake word detection system. The classifier integrates features from multiple sources including acoustic features, environmental information, lexical information, linguistic information, and contextual information to make holistic determinations about wake word presence, thereby improving detection accuracy while managing system complexity through integrated architecture

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The detection system uses composite feature sets that combine multiple types of information (acoustic, environmental, lexical, linguistic, and contextual features) similar to how composite materials combine different properties. This multi-faceted approach to feature composition enables the system to achieve higher detection accuracy by leveraging the strengths of different feature types while compensating for their individual weaknesses

Inventive Principle:
Principle #40Composite materials

2Measurement precision

If the detection system performs comprehensive analysis with multiple features and detectors, then the detection accuracy improves, but the processing latency increases

Engineering Contradiction:
Improvewake word detection accuracyVSAvoiddetection latency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary wake word detection using acoustic features and environmental information before committing to full speech processing. The classifier makes preliminary determinations about wake word presence based on available features, and only proceeds with more computationally intensive processing when the wake word is likely present, thereby reducing overall latency while maintaining accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The detection system is segmented into multiple independent detectors operating in series (acoustic feature detector, keyword spotter, speech activity detector, and classifier). Each detector processes features independently and contributes to the final decision, allowing for modular optimization where faster detectors can operate in parallel and slower detectors only process when needed, balancing accuracy and latency

Inventive Principle:
Principle #1Segmentation

3Reliability

If the system uses multiple detectors and comprehensive feature analysis, then false detection rates decrease, but the device complexity increases

Engineering Contradiction:
Improvefalse detection rateVSAvoiddetection system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The classifier serves as an intermediary that integrates outputs from multiple detectors (acoustic feature detector, keyword spotter, speech activity detector) and combines them with environmental, lexical, and contextual information. This intermediary component harmonizes the decisions from different detectors and feature sources, reducing false detections through comprehensive analysis while managing complexity through a centralized integration point rather than requiring complex interactions between all components

Inventive Principle:
Principle #24Intermediary (Mediator)

4Reliability

If the detection system processes more feature types, then the reliability of wake word detection improves, but the computational resources required increase

Engineering Contradiction:
Improvewake word detection reliabilityVSAvoidcomputational energy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system computes only the necessary subset of features required for reliable wake word detection rather than processing all possible feature types continuously. The classifier selectively processes acoustic, environmental, lexical, linguistic, and contextual features based on their relevance and availability, performing partial computation that achieves sufficient reliability without the excessive energy consumption of comprehensive processing of all possible features

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11657804B2Wake word detection modeling
Publication Date: 2023.05.23 AMAZON TECH INC
  • US11657804B2 patent drawing
  • US11657804B2 patent drawing
  • US11657804B2 patent drawing

AI summary

Features are disclosed for detecting words in audio using contextual information in addition to automatic speech recognition results. A detection model can be generated and used to determine whether a particular word, such as a keyword or “wake word,” has been uttered. The detection model can operate on features derived from an audio signal, contextual information associated with generation of the audio signal, and the like. In some embodiments, the detection model can be customized for particular users or groups of users based usage patterns associated with the users.