Voice Signal Clustering for Speaker Attribute Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speaker attribute estimation models require a large amount of training data with correct speaker attributes, making it costly and leading to decreased estimation accuracy due to overfitting when sufficient data is not available.

Innovation Solution

An estimation apparatus that clusters voice signals and uses a speaker attribute estimation model to identify clusters and estimate attributes based on feature similarities, allowing for accurate estimation even with insufficient training data by leveraging the most common attributes within clusters.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a large amount of training data with correct speaker attributes is used to construct a robust estimation model, then the reliability of speaker attribute estimation is improved, but the cost of data preparation increases and the device complexity increases

Engineering Contradiction:
Improverobust estimation capabilityVSAvoiddata preparation complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the training data into two distinct parts: voice data with speaker attributes (first voice data) and voice data without speaker attributes (second voice data). This segmentation allows the model to learn different aspects of speaker characteristics from each data type, achieving robust estimation without requiring all training data to be manually annotated with speaker attributes, thereby reducing data preparation complexity while maintaining reliability.

Inventive Principle:
Principle #1Segmentation

2Ease of manufacture

If a sufficient amount of training data with speaker attributes is not available, then the cost of data preparation is reduced, but the measurement precision of speaker attribute estimation decreases due to overfitting

Engineering Contradiction:
Improvedata preparation easeVSAvoidestimation accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary component - the speaker verification model trained on second voice data without speaker attributes. This intermediary model learns general speaker characteristics and voice features that serve as a foundation for the speaker attribute estimation model. By using this intermediary training step, the system achieves better estimation accuracy without requiring sufficient manually annotated training data, thus avoiding overfitting while maintaining ease of data preparation.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Device complexity

If only a speaker attribute estimation model trained on limited annotated data is used, then the device complexity is reduced, but the reliability of estimation for unknown speakers decreases

Engineering Contradiction:
Improvemodel structure simplicityVSAvoidestimation reliability for unknown speakers
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent applies preliminary action by training a speaker verification model on second voice data without speaker attributes before training the speaker attribute estimation model. This preliminary training step pre-processes the model to learn general speaker verification capabilities and voice features, which then serve as a strong foundation for the subsequent attribute estimation task. This two-stage preliminary preparation enables the model to generalize better to unknown speakers while keeping the overall device complexity manageable.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11996086B2Estimation device, estimation method, and estimation program
Publication Date: 2024.05.28 NIPPON TELEGRAPH & TELEPHONE CORP
  • US11996086B2 patent drawing
  • US11996086B2 patent drawing
  • US11996086B2 patent drawing

AI summary

An estimation apparatus clusters a group of voice signals including a voice signal having a speaker attribute to be estimated into a plurality of clusters. Subsequently, the estimation apparatus identifies, from the plurality of clusters, a cluster to which the voice signal to be estimated belongs. Next, the estimation apparatus uses a speaker attribute estimation model to estimate speaker attributes of respective voice signals in the identified cluster. After that, the estimation apparatus estimates an attribute of the entire cluster, by using an estimation result of the speaker attributes of the voice signals in the identified cluster, and outputs an estimation result of the speaker attribute of the entire cluster, as an estimation result of the speaker attribute of the voice signal to be estimated.