Speaker-Independent Keyword Model for Multi-User Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional electronic devices with speech recognition capabilities are limited by speaker-dependent keyword models, which are not effective in recognizing keywords spoken by multiple users due to differences in voice characteristics and pronunciation.

Innovation Solution

The development of a method and apparatus for obtaining a speaker-independent keyword model, where an electronic device generates a speaker-dependent model based on user samples and requests a speaker-independent model from a server, enabling accurate detection of keywords across various users.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a speaker-dependent keyword model is used, then the keyword detection accuracy for a specific user is improved, but the ability to recognize keywords spoken by multiple users deteriorates

Engineering Contradiction:
Improvekeyword detection accuracyVSAvoidmulti-user recognition capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a speaker-independent keyword model that serves multiple users simultaneously. The model is trained on keyword samples from multiple speakers and is designed to recognize keywords regardless of which speaker utters them, making the system universal across different users while maintaining detection accuracy

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent combines keyword samples from multiple different speakers into a single training dataset. By merging samples from various users and training a unified model on this combined data, the system achieves both multi-user adaptability and maintained detection accuracy through aggregated learning

Inventive Principle:
Principle #5Merging (Combining)

2Adaptability or versatility

If a speaker-independent keyword model is developed, then the multi-user recognition capability is improved, but the training data requirements and model complexity increase

Engineering Contradiction:
Improvemulti-user recognition capabilityVSAvoidmodel training complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent collects and prepares keyword samples from multiple speakers in advance during a training phase. By performing this data collection and model training beforehand, the system eliminates the need for real-time adaptation when new users speak, reducing operational complexity while maintaining multi-user capability

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically trains the speaker-independent keyword model using collected samples from multiple users without requiring manual configuration or intervention. The model self-adapts to multiple speakers through automated training processes, reducing the burden on users and simplifying deployment

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP3195310B1Keyword detection using speaker-independent keyword models for user-designated keywords
Publication Date: 2020.02.19 QUALCOMM INC
  • EP3195310B1 patent drawingFigure 1
  • EP3195310B1 patent drawingFigure 2
  • EP3195310B1 patent drawingFigure 3

AI summary

A method, which is performed by an electronic device, for obtaining a speaker-independent keyword model of a keyword designated by a user is disclosed. The method may include receiving at least one sample sound from the user indicative of the keyword. The method may also generate a speaker-dependent keyword model for the keyword based on the at least one sample sound, send a request for the speaker-independent keyword model of the keyword to a server in response to generating the speaker-dependent keyword model, and receive the speaker-independent keyword model adapted for detecting the keyword spoken by a plurality of users from the server.