Speaker Embedding Using Utterance Duration for Rhythm Capture

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speaker embedding techniques struggle to capture features such as utterance rhythm, leading to suboptimal performance in voice processing tasks.

Innovation Solution

A speaker embedding system that utilizes utterance unit segmentation information, specifically duration lengths for each utterance, to train a speaker identification model, enabling extraction of a speaker vector that captures utterance rhythm.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speaker embedding techniques using acoustic feature amounts are used, then speaker identification can be performed, but utterance rhythm features cannot be captured

Engineering Contradiction:
Improveutterance rhythm captureVSAvoidutterance rhythm information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent segments voice data into multiple utterance units and extracts duration information for each unit. This segmentation approach allows the system to capture temporal characteristics (utterance rhythm) that were previously lost in conventional acoustic feature extraction methods by analyzing each segment's duration individually and aggregating these temporal patterns.

Inventive Principle:
Principle #1Segmentation

2Reliability

If acoustic feature amounts are used for training, then speaker identification model can be trained, but performance in capturing speaking characteristics is insufficient

Engineering Contradiction:
Improvespeaker identification performanceVSAvoidutterance rhythm representation
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent introduces a new dimensional aspect to speaker embedding by incorporating duration information of utterance units. This adds a temporal dimension to the traditional acoustic feature space, enabling the model to capture both spectral characteristics and temporal rhythm patterns, thereby improving overall speaker identification performance and adaptability to different speaking styles.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12548573B2Speaker embedding device, speaker embedding method, and speaker embedding program
Publication Date: 2026.02.10 NT T INC
  • US12548573B2 patent drawing
  • US12548573B2 patent drawing
  • US12548573B2 patent drawing

AI summary

A speaker embedding apparatus includes processing circuitry configured to accept input of voice data, generate utterance unit segmentation information indicating a duration length for each utterance of a speaker in the input voice data, and use a duration length for each utterance indicated in the generated utterance unit segmentation information as training data and train a speaker identification model for outputting an identification result of a speaker when a duration length for each utterance of the speaker is input.