Speaker Recognition Model Update via Guided Voice Input

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional sentence-independent speaker recognition methods require users to utter a large number of sentences, leading to user fatigue and decreased accuracy due to ambient noise and quality deterioration.

Innovation Solution

An electronic apparatus and method that allows for the generation and updating of a speaker recognition model without requiring users to speak extensively, using a processor to guide user utterance and compare voice features with stored models, enabling efficient completion of the recognition process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the sentence-independent method is used to register a speaker recognition model, then the speaker recognition accuracy is improved, but the user convenience deteriorates due to the large amount of utterances required

Engineering Contradiction:
Improvespeaker recognition accuracyVSAvoiduser convenience
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system collects voice data from users during their normal use of the device before formal speaker recognition model registration is needed. This preliminary data collection occurs automatically as users interact with the device, building up a database of voice characteristics without requiring dedicated registration time from the user. When registration is eventually needed, the pre-collected data is already available to contribute to the speaker recognition model.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system performs automatic speaker recognition model registration without requiring active user participation in the traditional sense. The device autonomously collects voice data, processes it, and updates the speaker recognition model in the background during normal operation, eliminating the need for users to deliberately engage in lengthy registration procedures while still improving recognition accuracy.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If the recording time is increased to collect more utterance data, then the speaker recognition accuracy is improved, but the quality of the registered model deteriorates due to user fatigue and ambient noise

Engineering Contradiction:
Improvespeaker recognition accuracyVSAvoidmodel quality
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

Instead of continuous long-term recording that causes user fatigue, the system implements periodic, brief data collection intervals during which voice data is gathered. These periodic collection moments are distributed throughout normal device usage, allowing the system to accumulate sufficient data over time without requiring the user to engage in prolonged continuous recording sessions that would lead to fatigue and degraded quality.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The system performs preliminary data collection during casual, low-stakes moments of device usage before formal registration is needed. This approach gathers voice samples when users are naturally engaged with the device in relaxed states, rather than during dedicated registration sessions where users may be fatigued or overly conscious of their performance, thereby maintaining higher data quality.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3573054B1Electronic apparatus, controlling method and computer readable medium
Publication Date: 2023.12.13 SAMSUNG ELECTRONICS CO LTD
  • EP3573054B1 patent drawingFigure 1
  • EP3573054B1 patent drawingFigure 2A~2B
  • EP3573054B1 patent drawingFigure 3

AI summary

An electronic apparatus is provided. The electronic apparatus includes an inputter comprising input circuitry, a voice receiver comprising voice receiving circuitry, a storage, and a processor configured to: provide a guide prompting a user utterance based on user authentication being performed according to user information input through the inputter, generate a speaker recognition model corresponding to the user information based on a voice corresponding to the guide being received through the voice receiver, store the speaker recognition model in the storage, and identify a user corresponding to a voice received through the voice receiver based on the speaker recognition model updated by comparing a voice received through the voice receiver with the speaker recognition model.