Acoustic Model Reuse for Wakeword Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems require significant computing resources for full ASR and NLU processing, making them computationally expensive, and often necessitate a distributed computing environment to function efficiently.

Innovation Solution

The system re-uses a first acoustic model configured for limited ASR processing for wakeword detection, allowing multiple wakewords to be detected using different models, such as hidden Markov models, and associates these wakewords with different modes of operation, thereby eliminating the need for a second dedicated acoustic model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If full ASR and NLU processing is used for speech recognition, then accurate command recognition is achieved, but computational resources are significantly consumed

Engineering Contradiction:
Improvecommand recognition accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The speech processing system is segmented into two distinct paths: a lightweight wakeword detection path using acoustic models for simple keyword recognition, and a full ASR/NLU path for complex command processing. This segmentation allows the system to use minimal computational resources for routine wake-up detection while reserving full processing power for actual command execution.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Instead of applying full ASR and NLU processing to all audio inputs, the system applies partial processing only to audio segments that contain wakewords. The majority of audio inputs that do not trigger wakewords are discarded without expensive processing, significantly reducing overall computational resource consumption.

Inventive Principle:
Principle #16Partial or excessive action

2Measurement precision

If a dedicated second acoustic model is created for wakeword detection, then wakeword detection accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvewakeword detection accuracyVSAvoidnumber of acoustic models
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The existing acoustic models designed for full ASR processing are made multi-functional by reusing them for wakeword detection. The same acoustic models that process complete speech commands are also employed to detect wakewords, eliminating the need for separate dedicated wakeword detection models and reducing overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11308939B1Wakeword detection using multi-word model
Publication Date: 2022.04.19 AMAZON TECH INC
  • US11308939B1 patent drawing
  • US11308939B1 patent drawing
  • US11308939B1 patent drawing

AI summary

A system and method performs wakeword detection and automatic speech recognition using the same acoustic model. A mapping engine maps phones/senones output by the acoustic model to phones/senones corresponding to the wakeword. A hidden Markov model (HMM) may determine that the wakeword is present in audio data; the HMM may have multiple paths for multiple wakewords or may have multiple models. Once the wakeword is detected, ASR is performed using the acoustic model.