Speaker Recognition Threshold Adjustment via Context Similarity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Text-independent speaker recognition systems face challenges in usability and performance compared to text-dependent methods, with low usability in text-independent systems due to the need for specific utterances and lower recognition accuracy when verified utterances differ from registered ones.
Innovation Solution
An electronic device that stores a speaker model with acoustic features and context information, adjusting the threshold value for authentication based on the similarity between context information of registered and verified user voices, allowing for flexible voice recognition without requiring exact match utterances.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If text-independent speaker recognition method is used, then usability is improved (users can freely utter any voice), but speaker recognition accuracy deteriorates (low performance compared to text-dependent method)
Solution Approach 1:
The patent dynamically adjusts the threshold parameter for speaker verification based on context information similarity. When context information between registered and verified voices matches well, the threshold is raised to ensure accuracy; when context information differs significantly, the threshold is lowered to maintain usability. This parameter adaptation resolves the contradiction by making the system both flexible (text-independent) and accurate contextually.
2Measurement precision
If text-dependent speaker recognition method is used, then speaker recognition accuracy is improved (high accuracy), but usability deteriorates (users must use specific utterance words)
Solution Approach 1:
The patent introduces dynamic threshold adjustment based on context information similarity, transforming the static verification process into a dynamic one. The system adapts its strictness level in real-time: being more stringent when contexts match (ensuring accuracy) and more lenient when contexts differ (improving usability). This dynamic behavior allows the system to achieve high accuracy without requiring specific utterance words.
3Device complexity
If fixed threshold value is used for authentication, then system complexity is reduced (simple comparison), but recognition accuracy deteriorates (cannot adapt to different context scenarios)
Solution Approach 1:
The patent changes the threshold parameter dynamically based on context information similarity calculations. Instead of using a fixed threshold, the system computes a similarity metric between context information of registered and verified voices, then adjusts the threshold accordingly. This adds adaptability without significantly increasing system complexity, as the adjustment is based on a straightforward similarity comparison.
Data Source
AI summary
The present disclosure provides an electronic device and a control method thereof. The electronic device of the present disclosure includes: a memory in which a speaker model including acoustic characteristics and context information of a first user voice is stored; and a processor for comparing a degree of similarity between the acoustic characteristics of the first user included in the speaker model and the acoustic characteristics of a second user voice, with a threshold value changing according to a degree of similarity between the context information included in the speaker model and the context information of the second user voice, and then performing authentication on the second user voice.


