Speaker Identification Using Segmented Wake-Word Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in accurately approximating presence ground truth and authenticating user profiles in dynamic environments, such as homes, businesses, and vehicles, using electronic devices like voice interface devices and touch interface devices.

Innovation Solution

The system employs speaker identification processes, wake-word detection, and authentication-based processes to determine user presence and authenticate user profiles. It uses audio data from microphones to perform speaker identification, detects wake words to initiate processes, and applies authentication criteria to verify user identity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speaker identification and authentication processes are implemented to accurately determine user presence and verify user profiles, then authentication accuracy and security are improved, but system complexity and processing time increase

Engineering Contradiction:
Improveauthentication accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The authentication process is divided into multiple independent stages: wake-word detection, speaker identification, and user profile authentication. Each stage operates independently with its own processing logic and can be optimized separately, reducing overall system complexity while maintaining high authentication accuracy through sequential verification.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs wake-word detection and speaker identification before completing the full user profile authentication. This preliminary action filters out non-user interactions and identifies potential users early, reducing the complexity of subsequent authentication processes and improving overall system efficiency.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If multiple authentication processes are performed sequentially to verify user presence and identity, then authentication reliability is improved, but processing time and system response speed decrease

Engineering Contradiction:
Improveauthentication reliabilityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs wake-word detection and speaker identification as preliminary actions before completing user profile authentication. This early filtering reduces the number of full authentication cycles needed, improving reliability while reducing overall processing time by eliminating unnecessary authentication steps.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The authentication system uses feedback from each stage to determine whether to proceed to the next stage. If wake-word detection fails or speaker identification does not match, the system immediately terminates the authentication process, reducing processing time for unsuccessful attempts while maintaining high reliability for successful authentications.

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The system effectively approximates presence ground truth and authenticates user profiles with high accuracy, enabling improved user interaction with electronic devices and enhancing security in various environments.

Implementation Method 1

The system includes a microphone configured to capture audio signals

Methodology Applied
Scientific EffectAcoustic detection: Sound

Data Source

PatentUS12236957B1Authenticating a user profile with devices
Publication Date: 2025.02.25 AMAZON TECH INC
  • US12236957B1 patent drawing
  • US12236957B1 patent drawing
  • US12236957B1 patent drawing

AI summary

Systems and methods for presence ground truth approximation and utilization are disclosed. For example, a system detects the presence of a predefined subject, such as a person associated with a given user profile, and/or determines that authentication criteria for performing an action in association with the user profile has been satisfied. A period of time to associate data is determined, and data of one or more data types is labeled as being associated with the speaker identification event. That data may be formatted and input into one or more models to train those models to more accurately detect presence and/or determine whether authentication of a user profile should succeed.