Lip-password Speaker Verification via HMM Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing lip motion-based speaker verification systems face challenges in distinguishing between speakers due to insufficient feature representation and inadequate modeling of lip motion characteristics, particularly in noisy environments and with varying password lengths, leading to poor performance and susceptibility to impostors.
Innovation Solution
A multi-boosted Hidden Markov Model (HMM) approach is integrated with a random subspace method and data sharing scheme to segment and model lip-password sequences into distinguishable subunits, enhancing discriminative learning and verification accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional single Gaussian Mixture Model or single Hidden Markov Model is used for lip motion modeling, then the system is simple to implement, but the verification performance is insufficient and cannot differentiate similar lip motions between different speakers
Solution Approach 1:
The patent segments the lip-password sequence into multiple distinguishable subunits based on lip motion characteristics. Each subunit is modeled separately using HMM, allowing the system to capture temporal variations within each segment while maintaining overall sequence verification. This segmentation enables better differentiation of similar lip motions by analyzing local patterns rather than treating the entire sequence as a single unit.
Solution Approach 2:
The patent combines multiple HMM models into an ensemble verification system. Multiple HMMs are trained on different feature subsets or with different initializations, and their verification results are fused to make the final authentication decision. This merging of multiple models improves verification performance and robustness against impostors while maintaining computational feasibility through efficient fusion strategies.
2Reliability
If acoustic speech signals are used for speaker verification, then the system can verify speaker identity, but the performance degrades dramatically in noisy environments or with multiple talkers
Solution Approach 1:
The patent replaces acoustic signal processing with visual lip motion analysis. Instead of processing sound waves that are easily corrupted by background noise, the system processes video sequences of lip movements. This substitution of the sensing modality from acoustic to visual domain inherently provides robustness to acoustic interference, allowing verification to proceed reliably in noisy environments where acoustic-based systems would fail.
3Object-affected harmful factors
If lip motion features are used for speaker verification, then the system is insensitive to background noise, but the feature representation is insufficient to distinguish biometric properties between different speakers
Solution Approach 1:
The patent applies local quality by extracting and analyzing specific localized features from lip motion sequences that are most discriminative for speaker identification. Instead of treating all lip motion pixels equally, the system identifies and emphasizes key regions and temporal patterns that carry speaker-specific biometric information. This selective focus on locally discriminative features enhances the ability to distinguish between different speakers while maintaining noise insensitivity.
Solution Approach 2:
The patent enhances feature representation by incorporating temporal dimension analysis through HMM modeling. Rather than relying solely on spatial lip shape features from individual frames, the system models the temporal evolution of lip motions across sequences. This addition of the time dimension creates a more rich and discriminative feature space that captures dynamic speaker characteristics, enabling better differentiation between speakers who may have similar static lip shapes.
4Reliability
If a password protected biometric system is implemented using acoustic signals, then double security is achieved, but the password information is easily perceived and intercepted by listeners
Solution Approach 1:
The patent replaces acoustic password transmission with visual lip motion encoding. Instead of speakers uttering audible passwords that can be intercepted by microphones or human listeners, the system captures and verifies passwords through visual observation of lip movements. This substitution of the communication channel from acoustic to visual domain eliminates the risk of password interception through acoustic eavesdropping, while maintaining the double security architecture of password plus biometric verification.
Data Source
AI summary
A lip-based speaker verification system for identifying a speaker using a modality of lip motions; wherein an identification key of the speaker comprising one or more passwords; wherein the one or more passwords are embedded into lip motions of the speaker; wherein the speaker is verified by underlying dynamic characteristics of the lip motions; and wherein the speaker is required to match the one or more passwords embedded in the lip motions with registered information in a database. That is, in the case where the target speaker saying the wrong password or even in the case where an impostor knowing and saying the correct password, the nonconformities will be detected and the authentications/accesses will be denied.


