Voice Interpretation Device Detecting Synthesized Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice synthesis methods used for authentication in terminals are vulnerable to security threats as they cannot distinguish between actual and synthesized voices, allowing unauthorized access.
Innovation Solution
A voice interpretation device equipped with a processor and microphone that determines whether received audio is an actual voice or a synthesized voice using a difference model, outputting notifications to enhance security by rejecting synthesized voice authentication.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If voice synthesis is used for authentication, then authentication functionality is provided, but security vulnerability increases as synthesized voice cannot be distinguished from actual voice
Solution Approach 1:
The patent introduces a voiceprint analysis module as an intermediary that analyzes voice characteristics to determine whether the input voice is synthesized or actual. This mediator layer sits between the voice input and authentication decision, extracting biometric features that are difficult to replicate in synthesized voice, thereby maintaining security while preserving authentication functionality.
Solution Approach 2:
The patent replaces traditional text-based or simple voice recognition authentication with a more advanced voiceprint biometric authentication system. By substituting the authentication mechanism from text comparison to voiceprint analysis, the system gains the ability to detect synthesized voice through analysis of voice generation characteristics, thus improving security without sacrificing authentication capability.
2Ease of operation
If text-based authentication is used, then authentication is provided, but security is compromised as another person's voice can be used to unlock
Solution Approach 1:
The patent replaces text-based authentication with voiceprint biometric authentication. This substitution maintains ease of operation through voice-based input while significantly improving security by analyzing unique voice characteristics that are difficult to replicate, preventing unauthorized access through synthesized or imitated voices.
Solution Approach 2:
The patent changes the authentication parameter from text string matching to voiceprint biometric features. By analyzing parameters such as voice generation characteristics, spectral features, and temporal patterns that are inherent to human voice production, the system maintains user-friendly operation while preventing unauthorized authentication through synthesized voice.
3Ease of operation
If voice recognition is used for terminal control, then convenience is improved, but security vulnerability increases due to inability to detect synthesized voice
Solution Approach 1:
The patent introduces a voiceprint analysis module as an intermediary that sits between voice input and terminal control execution. This mediator analyzes voice characteristics to detect synthesized voice while allowing legitimate voice commands to proceed, thus maintaining convenience for authorized users while blocking unauthorized control through synthesized voice.
Solution Approach 2:
The patent replaces simple voice recognition with voiceprint biometric authentication for terminal control. This substitution maintains hands-free convenient operation while adding a security layer that can detect synthesized voice through analysis of voice generation characteristics, preventing unauthorized terminal control.
Data Source
AI summary
An apparatus that includes a microphone and a processor. The processor is configured to receive, via the microphone, audio comprising voice of a person, and determine whether the received audio is an actual voice or a synthesized voice. The apparatus also provides a first notification indicating that the received audio is the actual voice when the received audio is the actual voice, and provides a second notification indicating that the received audio is the synthesized voice when the received audio is the synthesized voice.


