Audio User Engagement Detection via Head Orientation Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing electronic devices struggle to accurately determine user engagement without relying on wake words, leading to inefficient processing of user speech inputs.

Innovation Solution

The system uses audio-based user engagement detection (UED) by extracting features from audio data to estimate head orientation, which serves as a proxy for user engagement, and combines this with image-based detection when a camera is available, allowing devices to determine if a user is engaged and process speech accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If audio-based user engagement detection is implemented, then processing efficiency is improved by avoiding unnecessary speech processing, but system complexity increases due to feature extraction and analysis components

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent replaces traditional wake-word-based mechanical triggering with audio-based engagement detection that analyzes acoustic features (spectral centroid, spectral rolloff, zero-crossing rate) to determine user engagement states. This substitution enables more efficient processing by continuously monitoring audio characteristics without requiring explicit wake words, thereby improving productivity while managing system complexity through algorithmic approaches.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes operational parameters by dynamically adjusting processing based on detected engagement states. When disengaged states are detected through audio feature analysis, the system reduces or suspends speech processing operations. This parameter change approach allows the system to optimize processing efficiency by adapting its operational intensity to actual user engagement levels.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If wake word-based processing is used, then system complexity is reduced, but processing efficiency deteriorates due to unnecessary processing of non-engaged speech inputs

Engineering Contradiction:
Improvesystem complexityVSAvoidprocessing efficiency
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent substitutes wake-word-based mechanical triggering with continuous audio-based engagement detection. By analyzing acoustic features such as spectral centroid, spectral rolloff, and zero-crossing rate, the system can distinguish between engaged and disengaged states without requiring explicit wake words. This substitution improves processing efficiency by avoiding unnecessary speech processing operations while maintaining manageable system complexity through well-established audio analysis techniques.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system implements self-service by automatically detecting user engagement states through audio analysis and autonomously adjusting its processing behavior. The engagement detection component continuously monitors audio features and determines when speech processing should be activated or deactivated, enabling the system to serve itself by optimizing its own operational efficiency without requiring external control mechanisms.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12531064B1Audio-based user engagement detection
Publication Date: 2026.01.20 AMAZON TECH INC
  • US12531064B1 patent drawing
  • US12531064B1 patent drawing
  • US12531064B1 patent drawing

AI summary

A system can operate a speech-controlled device to perform user engagement detection (UED) processing to detect when speech represented in audio data is directed to the device. For example, the device may extract audio features from the audio data and process these audio features using a classifier to estimate an orientation of the user's head, which may be used as a proxy for user engagement. Thus, if the head orientation is within an engagement zone (which varies based on distance to the user), the device may determine that the user is engaged with the device and perform language processing on input speech. In contrast, if the head orientation is outside of the engagement zone, the device may determine that the user is not engaged and ignore the input speech. To enable additional functionality, the classifier may optionally output a coarse estimate of the head orientation along with the UED determination.