Speaker Recognition Trigger Phrase Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice user interfaces with speaker recognition systems face challenges in ensuring adequate authentication for commands related to personal information, as the existing speaker recognition processes may not provide sufficient security.

Innovation Solution

A method and system that perform a speaker change detection process and fuse results from text-dependent and text-independent speaker recognition processes on the trigger phrase and preceding/following speech to enhance authentication, using a combination of speaker recognition processes to verify the identity of the user.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple speaker recognition processes are performed on trigger phrase and surrounding speech, then authentication reliability is improved, but system complexity increases

Engineering Contradiction:
Improveauthentication reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The authentication process is segmented into multiple independent speaker recognition processes: a text-dependent speaker recognition process on the trigger phrase, and a text-independent speaker recognition process on surrounding speech. Each process operates independently and contributes to the final authentication decision, allowing the system to achieve high reliability without requiring a single complex authentication mechanism.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system merges results from multiple speaker recognition processes by combining the output of text-dependent speaker recognition on the trigger phrase with text-independent speaker recognition on surrounding speech. This fusion of multiple recognition results creates a more robust authentication mechanism that overcomes the limitations of individual processes.

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If speaker change detection is performed continuously, then speaker recognition accuracy is improved, but processing time increases

Engineering Contradiction:
Improvespeaker recognition accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary speaker change detection on the trigger phrase before initiating the full speaker recognition process. This preliminary action identifies potential speaker boundaries in advance, allowing the system to focus computational resources only on relevant speech segments and avoid unnecessary processing of entire audio streams.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The speaker change detection and speaker recognition processes are applied selectively to specific local segments of speech rather than continuously to the entire audio stream. The system focuses analysis on the trigger phrase and immediately surrounding speech segments where speaker changes are most relevant to authentication, rather than processing all speech uniformly.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11437022B2Performing speaker change detection and speaker recognition on a trigger phrase
Publication Date: 2022.09.06 CIRRUS LOGIC INC
  • US11437022B2 patent drawing
  • US11437022B2 patent drawing
  • US11437022B2 patent drawing

AI summary

A method of speaker recognition comprises receiving an audio signal representing speech. A speaker change detection process is performed on the received audio signal. A trigger phrase detection process is also performed on the received audio signal. On detecting the trigger phrase in the received audio signal, a speaker recognition process is performed on the detected trigger phrase and on any speech preceding the detected trigger phrase and following an immediately preceding speaker change.