Speaker Recognition Model Update via Guided Voice Input
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional sentence-independent speaker recognition methods require users to utter a large number of sentences, leading to user fatigue and decreased accuracy due to ambient noise and quality deterioration.
Innovation Solution
An electronic apparatus and method that allows for the generation and updating of a speaker recognition model without requiring users to speak extensively, using a processor to guide user utterance and compare voice features with stored models, enabling efficient completion of the recognition process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the sentence-independent method is used to register a speaker recognition model, then the speaker recognition accuracy is improved, but the user convenience deteriorates due to the large amount of utterances required
Solution Approach 1:
The system collects voice data from users during their normal use of the device before formal speaker recognition model registration is needed. This preliminary data collection occurs automatically as users interact with the device, building up a database of voice characteristics without requiring dedicated registration time from the user. When registration is eventually needed, the pre-collected data is already available to contribute to the speaker recognition model.
Solution Approach 2:
The system performs automatic speaker recognition model registration without requiring active user participation in the traditional sense. The device autonomously collects voice data, processes it, and updates the speaker recognition model in the background during normal operation, eliminating the need for users to deliberately engage in lengthy registration procedures while still improving recognition accuracy.
2Measurement precision
If the recording time is increased to collect more utterance data, then the speaker recognition accuracy is improved, but the quality of the registered model deteriorates due to user fatigue and ambient noise
Solution Approach 1:
Instead of continuous long-term recording that causes user fatigue, the system implements periodic, brief data collection intervals during which voice data is gathered. These periodic collection moments are distributed throughout normal device usage, allowing the system to accumulate sufficient data over time without requiring the user to engage in prolonged continuous recording sessions that would lead to fatigue and degraded quality.
Solution Approach 2:
The system performs preliminary data collection during casual, low-stakes moments of device usage before formal registration is needed. This approach gathers voice samples when users are naturally engaged with the device in relaxed states, rather than during dedicated registration sessions where users may be fatigued or overly conscious of their performance, thereby maintaining higher data quality.
Data Source
Figure 1
Figure 2A~2B
Figure 3
AI summary
An electronic apparatus is provided. The electronic apparatus includes an inputter comprising input circuitry, a voice receiver comprising voice receiving circuitry, a storage, and a processor configured to: provide a guide prompting a user utterance based on user authentication being performed according to user information input through the inputter, generate a speaker recognition model corresponding to the user information based on a voice corresponding to the guide being received through the voice receiver, store the speaker recognition model in the storage, and identify a user corresponding to a voice received through the voice receiver based on the speaker recognition model updated by comparing a voice received through the voice receiver with the speaker recognition model.