Speaker Identification via Semantic Matching and Feature Similarity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional technologies fail to identify a certain utterer when the utterance content does not match the preliminarily registered utterance content, and they are limited by relying on fixed keyword sets, which prevents accurate identification of utterers when they utter keywords other than the fixed ones.

Innovation Solution

An utterer identification method that performs voice recognition on input utterance data, selects the registered utterance content closest to the recognized content, and calculates similarity with feature quantities stored in associated databases to identify the utterer, even when the utterance content is not identical to the registered content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional speech recognition technology is used to identify utterers by comparing input patterns with registered patterns, then identification can be performed for exact matches, but identification fails when the utterance content does not match the registered content

Engineering Contradiction:
Improveutterer identification accuracyVSAvoidflexibility to handle varied utterance content
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent segments the utterer identification process into two independent stages: (1) speech content recognition to extract semantic meaning, and (2) speaker feature extraction to identify the utterer. This segmentation allows the system to handle variations in utterance content while maintaining reliable speaker identification, as the speaker features are extracted independently from the specific content being spoken.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces speech content recognition as an intermediary component that bridges the input utterance and the speaker identification process. This intermediary extracts the semantic meaning of the utterance, allowing the system to select appropriate registered speakers even when the exact utterance content doesn't match, thereby improving both reliability and adaptability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If fixed keyword sets are used for speaker recognition, then the system can identify registered utterers for specific keywords, but it cannot identify utterers when they use keywords other than the fixed set

Engineering Contradiction:
Improvespeaker identification precisionVSAvoidcapability to recognize varied keywords
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal speaker identification system that works across multiple keywords and utterance contents. By extracting speaker features independently of specific keyword matching and using speech content recognition to bridge different utterances, the system achieves multi-functionality where a single registered speaker profile can be identified across various keywords and phrasings, not just fixed keyword sets.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240112682A1Speaker identification method, speaker identification device, and non-transitory computer readable recording medium
Publication Date: 2024.04.04 PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
  • US20240112682A1 patent drawing
  • US20240112682A1 patent drawing
  • US20240112682A1 patent drawing

AI summary

An utterer identification device executes: performing voice recognition from input utterance data; selecting, from among a plurality of registered utterance contents set in advance, a registered utterance content closest to a recognized utterance content indicated by a result of the voice recognition as a selected utterance content; selecting, from among a plurality of databases respectively associated with the registered utterance contents, a database associated with the selected utterance content; calculating a similarity between a feature quantity of the input utterance data and a feature quantity stored in the selected database; and identifying a certain utterer on the basis of the similarity, and outputting a result of the identification.