Voice Authentication Using Personator Model Updates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Fixed vocabulary speaker authentication is vulnerable to impersonation as a personator can record or mimic a specific word/sentence to succeed, while free utterance authentication is cumbersome due to the need for a large number of utterances to generate a speaker model.
Innovation Solution
An electronic apparatus performs additional speaker authentication using a processor to compare voice input data with both a speaker model and a personator model, determining similarities to verify the user's voice and update the personator model if authentication fails, thereby enhancing security and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If fixed vocabulary speaker authentication is used, then authentication process is simple and speaker model generation is easy, but security is compromised as personators can record or mimic specific words
Solution Approach 1:
The authentication process is divided into two independent stages: fixed vocabulary authentication for initial access and free utterance authentication for ongoing verification. This segmentation allows each stage to have optimized characteristics - simplicity in the first stage and security in the second stage - resolving the contradiction between ease of operation and reliability.
Solution Approach 2:
The system performs preliminary fixed vocabulary authentication to establish initial user identity before enabling voice recognition services. This preliminary action creates a baseline authentication that is simple to perform, while subsequent free utterance authentication during service usage provides enhanced security without requiring users to re-establish their speaker model.
2Reliability
If free utterance speaker authentication is used, then security is improved as more voice data is required, but device complexity increases due to multiple authentication processes
Solution Approach 1:
The voice recognition service itself serves dual purposes: it provides the intended functionality while simultaneously collecting voice data for free utterance authentication. This multi-functionality eliminates the need for separate authentication processes, as the service usage context is leveraged for security verification, reducing device complexity while maintaining high security standards.
Solution Approach 2:
The system automatically performs free utterance authentication during voice recognition service usage without requiring separate user actions. The voice data collected during normal service operation is automatically analyzed for authentication purposes, making the security enhancement transparent to users and avoiding additional operational complexity.
3Reliability
If free utterance speaker authentication is used, then security is improved, but the number of utterances required for speaker model generation increases
Solution Approach 1:
The system performs preliminary fixed vocabulary authentication to establish initial user identity before enabling voice recognition services. This preliminary action creates a baseline authentication that is simple to perform, while subsequent free utterance authentication during service usage provides enhanced security without requiring users to re-establish their speaker model.
Solution Approach 2:
The free utterance authentication process operates continuously during voice recognition service usage rather than requiring a separate setup phase. As users naturally interact with the device through voice commands during service usage, the system continuously collects and analyzes voice data for authentication purposes, eliminating dedicated model generation time and integrating security verification into the normal service flow.
Data Source
AI summary
An electronic apparatus and a method thereof are provided. The electronic apparatus according to an embodiment includes a memory to store a speaker model including characteristic information of a voice of a user corresponding to a specific word, and a processor configured to, based on a first input voice corresponding to the specific word, provide a voice recognition service through the electronic apparatus based on a first authentication performed based on data on the first input voice and the speaker model, and perform a second authentication based on data corresponding to the specific word among data of the second input voice that is input while the voice recognition service is provided and the speaker model.


