Speaker Recognition Threshold Adjustment via Context Similarity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Text-independent speaker recognition systems face challenges in usability and performance compared to text-dependent methods, with low usability in text-independent systems due to the need for specific utterances and lower recognition accuracy when verified utterances differ from registered ones.

Innovation Solution

An electronic device that stores a speaker model with acoustic features and context information, adjusting the threshold value for authentication based on the similarity between context information of registered and verified user voices, allowing for flexible voice recognition without requiring exact match utterances.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If text-independent speaker recognition method is used, then usability is improved (users can freely utter any voice), but speaker recognition accuracy deteriorates (low performance compared to text-dependent method)

Engineering Contradiction:
ImproveusabilityVSAvoidspeaker recognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent dynamically adjusts the threshold parameter for speaker verification based on context information similarity. When context information between registered and verified voices matches well, the threshold is raised to ensure accuracy; when context information differs significantly, the threshold is lowered to maintain usability. This parameter adaptation resolves the contradiction by making the system both flexible (text-independent) and accurate contextually.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If text-dependent speaker recognition method is used, then speaker recognition accuracy is improved (high accuracy), but usability deteriorates (users must use specific utterance words)

Engineering Contradiction:
Improvespeaker recognition accuracyVSAvoidusability
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent introduces dynamic threshold adjustment based on context information similarity, transforming the static verification process into a dynamic one. The system adapts its strictness level in real-time: being more stringent when contexts match (ensuring accuracy) and more lenient when contexts differ (improving usability). This dynamic behavior allows the system to achieve high accuracy without requiring specific utterance words.

Inventive Principle:
Principle #15Dynamics

3Device complexity

If fixed threshold value is used for authentication, then system complexity is reduced (simple comparison), but recognition accuracy deteriorates (cannot adapt to different context scenarios)

Engineering Contradiction:
Improvesystem complexityVSAvoidrecognition accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent changes the threshold parameter dynamically based on context information similarity calculations. Instead of using a fixed threshold, the system computes a similarity metric between context information of registered and verified voices, then adjusts the threshold accordingly. This adds adaptability without significantly increasing system complexity, as the adjustment is based on a straightforward similarity comparison.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12002475B2Electronic device performing speaker recognition and control method thereof
Publication Date: 2024.06.04 SAMSUNG ELECTRONICS CO LTD
  • US12002475B2 patent drawing
  • US12002475B2 patent drawing
  • US12002475B2 patent drawing

AI summary

The present disclosure provides an electronic device and a control method thereof. The electronic device of the present disclosure includes: a memory in which a speaker model including acoustic characteristics and context information of a first user voice is stored; and a processor for comparing a degree of similarity between the acoustic characteristics of the first user included in the speaker model and the acoustic characteristics of a second user voice, with a threshold value changing according to a degree of similarity between the context information included in the speaker model and the context information of the second user voice, and then performing authentication on the second user voice.