Unobtrusive Speaker Verification via Background Voice Sampling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Text-independent speaker verification systems require a burdensome and intensive training process, while text-dependent systems are limited in recognition capabilities, often leading to frustration due to extraneous sounds captured during enrollment.

Innovation Solution

An electronic device receives voice samples during normal interaction with voice command-and-control features, classifies them into voice type categories, groups samples into user sets, generates voice models, and requests user identity for labeling, enabling unobtrusive training for both text-dependent and text-independent speaker verification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If text-independent speaker verification systems are implemented, then recognition flexibility and versatility are improved, but training process duration and complexity increase significantly

Engineering Contradiction:
Improverecognition flexibilityVSAvoidtraining process duration
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary voice sample collection during normal device usage before actual verification is needed. Voice samples are gathered continuously as users interact with the device, preparing training data in advance so that when verification is required, the model is already trained or can be quickly trained with collected samples.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically collects and processes voice samples without requiring user intervention for training. The electronic device itself gathers voice data during normal operation, automatically classifies samples by voice type, and generates speaker models without needing users to complete separate enrollment procedures.

Inventive Principle:
Principle #25Self-service

2Quantity of substance

If voice samples are collected during normal interaction, then training data quantity increases, but system complexity increases

Engineering Contradiction:
Improvetraining data quantityVSAvoidsystem complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system segments voice samples into different voice type categories based on characteristics such as background noise levels, recording conditions, and speech patterns. This segmentation allows the system to organize and process large volumes of voice data in manageable groups, reducing the computational complexity of training while maintaining data quantity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The voice collection mechanism serves multiple functions: it collects training data for speaker verification, improves speech recognition accuracy, and adapts to different usage scenarios. By making the voice collection system multi-functional, the patent reduces overall system complexity while increasing training data quantity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of time

If automatic voice model generation is implemented, then user enrollment time is reduced, but measurement precision requirements increase

Engineering Contradiction:
Improveuser enrollment timeVSAvoidvoice sample classification precision
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The system implements feedback loops where collected voice samples are continuously analyzed and classified. The classification results feed back into the training process, allowing the system to refine its voice type categorization and improve speaker model accuracy automatically. This feedback mechanism enables automatic model generation while maintaining sufficient precision through iterative refinement.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10152974B2Unobtrusive training for speaker verification
Publication Date: 2018.12.11 SENSORY INC
  • US10152974B2 patent drawing
  • US10152974B2 patent drawing
  • US10152974B2 patent drawing

AI summary

Techniques for implementing unobtrusive training for speaker verification are provided. In one embodiment, an electronic device can receive a plurality of voice samples uttered by one or more users as they interact with a voice command-and-control feature of the electronic device and, for each voice sample, assign the voice sample to one of a plurality of voice type categories. The electronic device can further group the voice samples assigned to each voice type category into one or more user sets, where each user set comprises voice samples likely to have been uttered by a unique user. The electronic device can then, for each user set: (1) generate a voice model, (2) issue, to the unique user, a request to provide an identity or name, and (3) label the voice model with the identity or name provided by the unique user.