Speaker Recognition Trigger Phrase Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice user interfaces with speaker recognition systems face challenges in ensuring adequate authentication for commands related to personal information, as the existing speaker recognition processes may not provide sufficient security.
Innovation Solution
A method and system that perform a speaker change detection process and fuse results from text-dependent and text-independent speaker recognition processes on the trigger phrase and preceding/following speech to enhance authentication, using a combination of speaker recognition processes to verify the identity of the user.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple speaker recognition processes are performed on trigger phrase and surrounding speech, then authentication reliability is improved, but system complexity increases
Solution Approach 1:
The authentication process is segmented into multiple independent speaker recognition processes: a text-dependent speaker recognition process on the trigger phrase, and a text-independent speaker recognition process on surrounding speech. Each process operates independently and contributes to the final authentication decision, allowing the system to achieve high reliability without requiring a single complex authentication mechanism.
Solution Approach 2:
The system merges results from multiple speaker recognition processes by combining the output of text-dependent speaker recognition on the trigger phrase with text-independent speaker recognition on surrounding speech. This fusion of multiple recognition results creates a more robust authentication mechanism that overcomes the limitations of individual processes.
2Measurement precision
If speaker change detection is performed continuously, then speaker recognition accuracy is improved, but processing time increases
Solution Approach 1:
The system performs preliminary speaker change detection on the trigger phrase before initiating the full speaker recognition process. This preliminary action identifies potential speaker boundaries in advance, allowing the system to focus computational resources only on relevant speech segments and avoid unnecessary processing of entire audio streams.
Solution Approach 2:
The speaker change detection and speaker recognition processes are applied selectively to specific local segments of speech rather than continuously to the entire audio stream. The system focuses analysis on the trigger phrase and immediately surrounding speech segments where speaker changes are most relevant to authentication, rather than processing all speech uniformly.
Data Source
AI summary
A method of speaker recognition comprises receiving an audio signal representing speech. A speaker change detection process is performed on the received audio signal. A trigger phrase detection process is also performed on the received audio signal. On detecting the trigger phrase in the received audio signal, a speaker recognition process is performed on the detected trigger phrase and on any speech preceding the detected trigger phrase and following an immediately preceding speaker change.


