Passive Voice Biometric Identification in Secure Calls
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice biometric systems for secure environments, such as correctional facilities, are limited in their ability to continuously monitor and verify the identity of speakers throughout a call, require cumbersome enrollment processes, and fail to identify the receiving party, posing security risks due to the impracticality of monitoring all telephone traffic and the lack of identification in voicemail systems.
Innovation Solution
A passive detection system that automatically creates and verifies Biometric Voice Prints (BVPs) without formal enrollment, allowing for continuous identification of speakers in both directions of a call and alerting security personnel to known individuals or changes in speakers, using pre-recorded calls to generate high-quality voice prints within seconds of net speech, and integrating with voicemail systems for real-time authentication.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional voice biometric enrollment is used, then identity verification accuracy is improved, but the enrollment process becomes cumbersome and time-consuming
Solution Approach 1:
The system performs preliminary voice print extraction and enrollment during the call setup phase before the actual conversation begins. This allows the enrollment to be completed in advance, making the process seamless and not interfering with the user experience while still ensuring accurate identity verification when needed
Solution Approach 2:
The voice biometric system operates autonomously by automatically extracting voice prints from captured audio, comparing them against the database, and verifying identities without requiring manual intervention or user cooperation beyond the initial call setup, thereby simplifying the enrollment process while maintaining accuracy
2Loss of time
If verification is performed only at the beginning of interaction, then processing time is reduced, but identity security is compromised when speakers change
Solution Approach 1:
The system continuously monitors the call audio throughout the entire interaction, repeatedly extracting voice prints and comparing them against the enrolled database. This continuous verification ensures that any speaker changes are detected immediately, maintaining high reliability without requiring lengthy processing intervals
Solution Approach 2:
The system performs periodic voice print extraction and verification at regular intervals during the call, combining efficient batch processing with continuous security monitoring. This approach balances processing time requirements with the need to detect speaker changes, maintaining both speed and reliability
3Reliability
If all telephone traffic is monitored, then security detection capability is improved, but resource requirements and operational complexity increase
Solution Approach 1:
The system extracts only the relevant voice biometric features from the telephone audio traffic using automatic speech recognition and voice print extraction algorithms. By focusing on these specific acoustic features rather than analyzing entire conversations, the system achieves high security detection capability while minimizing resource consumption and operational complexity
Solution Approach 2:
The system replaces manual monitoring and verification processes with automated voice biometric analysis using machine learning algorithms and pattern recognition. This substitution of automated computational methods for manual operations significantly reduces the complexity of monitoring all telephone traffic while maintaining or improving security detection capability
Data Source
AI summary
A method for passive enrollment and identification of a telephone caller to a called telephone number, comprising the steps of audio recording a telephone call; identifying and separating any multiple speakers on the telephone call and specifying a one of the multiple speakers; creating a net speech portion of the telephone call by trimming portions of audio recording from the beginning and end of the audio recording; processing the net speech portion against an existing Biometric Voice Print (BVP) database; creating a new BVP for the at least one of the multiple speakers if no match of the net speech portion against the BVP database is found in the processing step; comparing subsequent calls against the BVP, whether existing or created, to identify the at least one of the multiple speakers; and associating in a cluster all subsequent calls having voice prints matching the BVP.


