Automatic Speaker Feature Generation for Text-Dependent Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text-dependent speaker verification (TD-SV) techniques for automated assistants are inadequate in verifying users with sufficient confidence, leading to prolonged interactions and resource utilization due to the need for additional authentication prompts when invocation phrases are absent or insufficient, and require burdensome enrollment procedures.
Innovation Solution
Generate speaker features for specific TD-SVs based on normal user interactions using audio data portions identified through speech recognition and authentication measures, without separate enrollment procedures, and authenticate users by comparing utterance features to these speaker features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If TD-SV is constrained to invocation phrases only, then the verification process is simple to implement, but verification confidence is insufficient leading to prolonged interactions
Solution Approach 1:
The patent segments the spoken utterance into multiple portions and applies different TD-SV models to each segment. Instead of relying on a single invocation phrase verification, the system divides the utterance and performs multiple verification passes, each contributing to the overall verification confidence. This segmentation allows the system to maintain simplicity while improving reliability through cumulative verification evidence.
Solution Approach 2:
The patent performs TD-SV verification on portions of the utterance that may exceed the minimal invocation phrase requirement. By verifying multiple segments including but not limited to the invocation phrase, the system gathers excessive verification evidence that strengthens confidence in the verification decision, reducing the need for prolonged interactions or additional authentication prompts.
2Measurement precision
If TD-SV requires explicit enrollment procedure, then speaker features can be generated accurately, but user interaction is prolonged and resources are overutilized
Solution Approach 1:
The patent performs speaker feature generation as a preliminary action during normal user interactions before explicit enrollment is required. By continuously processing and storing speaker features from routine utterances, the system prepares verification data in advance, eliminating the need for time-consuming enrollment procedures when authentication is needed.
Solution Approach 2:
The system performs self-service by automatically generating and updating speaker features from user's normal interactions without requiring explicit user enrollment actions. The TD-SV model continuously learns and adapts to the user's voice characteristics during regular usage, making the enrollment process unnecessary while maintaining high speaker feature accuracy.
3Reliability
If multiple TD-SVs are used for different terms, then verification accuracy is enhanced, but system complexity increases
Solution Approach 1:
The patent implements a universal TD-SV framework where a single verification system handles multiple terms and utterance portions. Instead of requiring completely separate verification systems for each term, the same TD-SV model is applied universally across different segments and terms, achieving high verification accuracy while controlling system complexity through reuse of the core verification mechanism.
Solution Approach 2:
The patent segments the verification process into multiple portions that are processed sequentially or in parallel, with each segment handled by the same TD-SV model. This segmentation approach allows the system to maintain high verification accuracy by examining multiple aspects of the utterance while avoiding the complexity of training and maintaining multiple distinct verification models.
Data Source
AI summary
Implementations relate to automatic generation of speaker features for each of one or more particular text-dependent speaker verifications (TD-SVs) for a user. Implementations can generate speaker features for a particular TD-SV using instances of audio data that each capture a corresponding spoken utterance of the user during normal non-enrollment interactions with an automated assistant via one or more respective assistant devices. For example, a portion of an instance of audio data can be used in response to: (a) determining that recognized term(s) for the spoken utterance captured by that the portion correspond to the particular TD-SV; and (b) determining that an authentication measure, for the user and for the spoken utterance, satisfies a threshold. Implementations additionally or alternatively relate to utilization of speaker features, for each of one or more particular TD-SVs for a user, in determining whether to authenticate a spoken utterance for the user.


