Voice One-Time Passphrase Authentication Against Man-in-the-Middle Fraud
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for verifying caller identity are cumbersome, insecure, and prone to man-in-the-middle attacks, allowing fraudsters to exploit vulnerabilities and capture sensitive information during telephone communications.
Innovation Solution
A voice-based one-time password (OTP) system that generates an OTP prompt based on contextual information, requiring the user to speak the OTP aloud, using speaker recognition, liveness detection, and fraud detection features to authenticate the user and operation request.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional caller identification services are used, then caller information can be transmitted to callees, but the system becomes vulnerable to spoofing attacks and man-in-the-middle attacks
Solution Approach 1:
The system performs preliminary authentication actions by generating and transmitting a voice OTP to the callee's device before the actual communication occurs. This advance verification step ensures that the caller's identity is confirmed through voice biometrics and liveness detection, preventing spoofing attacks from succeeding. The OTP is generated based on contextual information and sent through a secure channel, establishing trust before the vulnerable communication phase begins.
Solution Approach 2:
The system introduces a voice OTP intermediary mechanism that mediates between the caller and callee. Instead of directly trusting the caller identification provided by traditional services, the system uses a voice-based one-time password as an intermediate verification layer. This intermediary element (the voice OTP) acts as a trusted mediator that confirms the caller's identity through voice biometrics, preventing man-in-the-middle attacks by ensuring that only authenticated voices can proceed with communication.
2Reliability
If knowledge-based questions are used for authentication, then user identity can be verified, but the process becomes cumbersome and time-consuming
Solution Approach 1:
The system replaces the mechanical interaction of answering knowledge-based questions with a voice-based biometric authentication mechanism. Instead of requiring users to recall and articulate private information (mechanical cognitive process), the system captures and analyzes voice characteristics through acoustic processing. This substitution maintains high verification accuracy while significantly reducing the time and cognitive effort required, as voice authentication occurs automatically during natural speech rather than requiring deliberate recall of security questions.
Solution Approach 2:
The voice OTP system enables self-service authentication where the user's own voice characteristics serve as the authentication credential. The system automatically captures voice samples during the call, extracts biometric features, and performs verification without requiring the user to manually answer security questions or perform additional actions. This self-service approach maintains security while reducing authentication time, as the verification process is seamlessly integrated into the natural communication flow.
3Ease of operation
If the telephone communication channel is used for information exchange, then communication convenience is maintained, but the channel becomes increasingly untrustworthy due to spoofing and social engineering vulnerabilities
Solution Approach 1:
The system applies preliminary anti-action by proactively defending against spoofing and social engineering attacks before they can compromise the communication. The voice OTP mechanism pre-verifies the caller's identity through voice biometrics and liveness detection, creating a protective barrier that neutralizes potential threats. This preliminary security measure maintains communication convenience while counteracting the untrustworthiness of the telephone channel, as the authentication occurs automatically without disrupting the natural flow of conversation.
Solution Approach 2:
The system implements feedback by continuously monitoring voice characteristics during the authentication process and providing real-time verification results. The voice OTP system analyzes acoustic features, speaker recognition scores, and liveness indicators, then feeds this information back to determine authentication outcomes. This feedback mechanism maintains communication trustworthiness by dynamically assessing voice authenticity while preserving ease of operation, as the process occurs seamlessly during normal speech interactions without requiring user awareness or additional actions.
Data Source
AI summary
Embodiments described herein provide for automatically authenticating operation requests and end-users who submit operation requests during contact events. A server obtains an operation request for an operation originated at an end-user device. The server generates a voice-based one-time password (OTP) using contextual information associated with the requested operation. The server generates and transmits an OTP prompt having text representing the OTP for display at a user interface of the user device. The server receives a response including an audio signal that contains the recording of the user speaking the OTP text aloud. The server uses the audio signal to authenticate the user and the operation request based on the speaker's voice, the accuracy of the user speaking the OTP, and liveness or fraud detection features extracted from the audio signal or metadata from the user device.


