Speaker Verification Model Training via Sensor-Gated Voice Capture

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice recognition systems face challenges in simplifying user registration procedures for voice recognition services while maintaining high accuracy and minimizing user inconvenience.

Innovation Solution

The system trains a speaker verification model by securely acquiring and updating voice data from registered users, simplifying the voice input process and enhancing verification accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice data is collected during user registration for speaker verification, then verification accuracy is improved, but user inconvenience and registration complexity increase

Engineering Contradiction:
Improvespeaker verification accuracyVSAvoiduser registration convenience
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system collects and stores voice data from multiple users during an initial registration phase before actual verification is needed. By performing this data collection action in advance, the system builds a comprehensive voiceprint database that enables accurate speaker verification without requiring users to provide additional voice samples during each verification event, thus eliminating the need for repeated registration procedures.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates voiceprint copies (feature extractions) from original voice recordings during registration. These copied representations are stored and used for verification, allowing the system to accurately identify speakers without requiring them to speak again during verification. The copying process transforms complex voice data into compact, comparable features that maintain verification accuracy while simplifying the user experience.

Inventive Principle:
Principle #26Copying

2Reliability

If voice data is securely stored and updated for trained speaker verification, then verification reliability is improved, but system complexity increases

Engineering Contradiction:
Improvespeaker verification reliabilityVSAvoidvoice data management complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system merges voice data from multiple users into a unified speaker verification model. By combining and consolidating voiceprints from different users during the registration phase, the system creates a comprehensive database that improves verification reliability across all users. This merging approach allows the system to maintain high reliability without requiring separate complex management systems for each user's data.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system transforms raw voice data into standardized voiceprint parameters through feature extraction and processing. By changing the parameters from raw audio signals to structured voiceprint representations, the system enables reliable speaker verification while simplifying data storage and management. This parameter transformation reduces complexity by converting unstructured voice data into comparable, standardized features that are easier to manage and process.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4195202B1Device for learning speaker authentication of registered user for voice recognition service, and method for operating same
Publication Date: 2024.11.13 SAMSUNG ELECTRONICS CO LTD
  • EP4195202B1 patent drawingFigure 1
  • EP4195202B1 patent drawingFigure 2
  • EP4195202B1 patent drawingFigure 3

AI summary

An electronic device according to various embodiments of the disclosure may include: a microphone, at least one sensor, a memory storing a speaker verification model for verifying a voice of a registered user, and at least one processor operatively connected to the microphone, the at least one sensor, and the memory, wherein the at least one processor is configured to: identify whether user verification information is received through the at least one sensor within a designated time interval before or after a time point at which an uttered voice is received based on receiving the uttered voice through the microphone, and updates the speaker verification model using the uttered voice based on the user verification being completed according to the user verification information.