Speaker Recognition Anchor Model Generation for Embedded Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speaker recognition systems require a large number of anchors to accurately describe a target speaker, leading to increased computation load and difficulty in implementation on embedded systems, as they can only reflect the distribution of the anchor space.

Innovation Solution

The proposed solution involves generating a reference anchor set using enrollment speeches based on an anchor space, which includes principal and associate anchors, allowing for the creation of smaller-sized anchor models through speaker adaptation techniques, thereby reducing computational load and memory requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a large number of anchors are used to describe the target speaker better, then the speaker recognition accuracy is improved, but the computation load increases and it becomes more difficult to implement on embedded systems

Engineering Contradiction:
Improvespeaker recognition accuracyVSAvoidcomputation load
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential and representative anchors from the anchor space to form a compact reference anchor set. Instead of using all anchors or a large number of them, the system selectively extracts the most relevant anchors that capture the essential speaker characteristics, thereby reducing computation load while maintaining recognition accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the large anchor space into multiple clusters and selects representative anchors from each cluster. This segmentation approach allows the system to cover the entire speaker characteristic space with a smaller number of strategically selected anchors, reducing the overall reference anchor set size while preserving recognition accuracy.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If a large number of anchors are used to describe the target speaker better, then the speaker recognition accuracy is improved, but the memory requirements increase

Engineering Contradiction:
Improvespeaker recognition accuracyVSAvoidmemory requirements
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential and representative anchors from the anchor space to form a compact reference anchor set. Instead of using all anchors or a large number of them, the system selectively extracts the most relevant anchors that capture the essential speaker characteristics, thereby reducing computation load while maintaining recognition accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the large anchor space into multiple clusters and selects representative anchors from each cluster. This segmentation approach allows the system to cover the entire speaker characteristic space with a smaller number of strategically selected anchors, reducing the overall reference anchor set size while preserving recognition accuracy.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9595260B2Modeling device and method for speaker recognition, and speaker recognition system
Publication Date: 2017.03.14 PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
  • US9595260B2 patent drawing
  • US9595260B2 patent drawing
  • US9595260B2 patent drawing

AI summary

A modeling device comprises a front end which receives enrollment speech data from each target speaker, a reference anchor set generation unit which generates a reference anchor set using the enrollment speech data based on an anchor space, and a voice print generation unit which generates voice prints based on the reference anchor set and the enrollment speech data. By taking the enrollment speech and speaker adaptation technique into account, anchor models with a smaller size can be generated, so reliable and robust speaker recognition with a smaller size reference anchor set is possible.