Speaker Verification Using Global Model and Noise Compensation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speaker verification systems face inefficiencies in large-scale applications due to high false acceptance and rejection rates, long enrolling sessions, and the need for extensive verification data, which hinder usability and increase traffic load, especially in telephone network environments.
Innovation Solution
A biometric control method using voice verification with limited enrolling data and short verification sentences, employing an equalization method and unsupervised intra-speaker and noise compensation technique, allowing simultaneous authentication across multiple users and reducing the need for additional infrastructure, while maintaining low error rates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speaker verification systems use traditional methods with multiple client templates, then verification accuracy is improved, but system complexity and processing time increase significantly
Solution Approach 1:
The patent extracts and removes the client-specific template matching step from the verification process. Instead of comparing verification speech against multiple client templates, the system only maintains a single global speaker model and verifies whether the verification speech belongs to any speaker in the database, eliminating the need for n identification distance evaluations.
Solution Approach 2:
The patent implements a universal speaker verification system with a single global speaker model that serves all clients. This universal model replaces the need for individual client-specific templates, allowing the system to verify any speaker without requiring separate model maintenance and comparison for each client.
2Reliability
If speaker verification systems require long enrolling sessions and extensive verification data, then error rates are reduced, but usability and traffic load are negatively impacted
Solution Approach 1:
The patent changes the fundamental parameter of data requirement from extensive to limited. By using a global speaker model with unsupervised adaptation that automatically adjusts to individual speaker characteristics during verification, the system achieves reliable error rates with minimal enrolling data and short verification sentences.
Solution Approach 2:
The system performs unsupervised adaptation automatically during the verification process without requiring manual intervention or extensive pre-enrollment. The global speaker model self-adjusts to individual speaker characteristics on-the-fly, enabling the system to maintain low error rates with minimal user input and short verification sentences.
3Measurement precision
If speaker verification systems process multiple client templates, then comprehensive verification is achieved, but processing time and traffic load increase
Solution Approach 1:
The patent removes the computationally intensive step of comparing verification speech against multiple client templates. The system only performs single-model verification against the global speaker model, dramatically reducing processing time while maintaining verification coverage through unsupervised adaptation to individual speaker characteristics.
4Adaptability or versatility
If speaker verification systems implement unsupervised adaptation to capture voice variations, then adaptability is improved, but error propagation risk increases
Solution Approach 1:
The patent introduces a global speaker model as an intermediary between the verification speech and the decision process. This intermediary model performs unsupervised adaptation to capture voice variations while providing a stable reference framework that prevents error propagation, as adaptation errors are localized to the model adjustment rather than propagated through multiple template comparisons.
Data Source
AI summary
A large-scale attendance, productivity, activity and availability biometric control method using the telephone network, for individual client users with speaker verification technology based on limited enrolling data and short verification sentences. The method includes the steps of registering and enrolling a client user; generating, storing and indexing a template and a reference average spectrum with a client user PIN; prompting the client user during a future verification event to pronounce the enrolling/verification sentence associated with the PIN to provide a speech signal; estimating a verification distance between the PIN indexed template and the speech signal pronounced by the client user, using the PIN reference average spectrum indexed; validating the telephone number; deciding if the pronounced speech signal pronounced by the client user has been validated based on the verification distance, with unsupervised compensation of the noisy input signal's spectrum if its difference from the speaker model is small, instead of adapting the user model spectrum; optionally repeating the steps of prompting, estimating, validating and deciding for a limited number of times; and accepting or rejecting the future verification event.


