Speaker Co-occurrence Model for Multi-Speaker Voice Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice data analyzing devices fail to consider the relationships among speakers, leading to decreased recognition accuracy, particularly in scenarios involving multiple speakers, such as criminal investigations or phishing scams.

Innovation Solution

A voice data analyzing device that derives speaker models and co-occurrence models from segmented voice data, incorporating the relationships between speakers through speaker model learning and co-occurrence model derivation, enabling accurate recognition of speakers in conversations involving multiple individuals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speaker models are learned independently for each speaker using voice data and speaker labels, then the speaker recognition process can be executed for each speaker model independently, but the relationship between speakers cannot be utilized leading to deteriorated recognition accuracy

Engineering Contradiction:
Improveindependent speaker model learningVSAvoidspeaker recognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent merges independent speaker model learning with relationship analysis by introducing a speaker relationship model that captures co-occurrence patterns between speakers. The system combines speaker-specific acoustic models with a relationship layer that models how speakers interact and co-appear in conversations, thereby utilizing relationship information to improve recognition accuracy while maintaining the modular structure of independent speaker model learning.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent adds another dimension to the traditional speaker recognition framework by incorporating speaker relationship information as a separate layer. Instead of only modeling speakers independently in the acoustic feature space, the system extends the representation to include relationship features that capture co-occurrence patterns, effectively adding a relational dimension to the recognition problem.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If clustering of learned speakers is conducted by determining vocal tract length expansion/contraction coefficient, then speakers can be grouped based on physical characteristics, but the relationship between speakers is still not discussed leading to limited recognition capability

Engineering Contradiction:
Improvespeaker clustering capabilityVSAvoidspeaker relationship recognition
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent changes the parameters used for speaker representation by introducing relationship-based features alongside traditional acoustic parameters. Instead of relying solely on vocal tract length coefficients for clustering, the system incorporates speaker relationship models that capture co-occurrence patterns, thereby changing the parameter space to include relational information that improves recognition capability.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8954327B2Voice data analyzing device, voice data analyzing method, and voice data analyzing program
Publication Date: 2015.02.10 NEC CORP
  • US8954327B2 patent drawing
  • US8954327B2 patent drawing
  • US8954327B2 patent drawing

AI summary

A voice data analyzing device comprises speaker model deriving means which derives speaker models as models each specifying character of voice of each speaker from voice data including a plurality of utterances to each of which a speaker label as information for identifying a speaker has been assigned and speaker co-occurrence model deriving means which derives a speaker co-occurrence model as a model representing the strength of co-occurrence relationship among the speakers from session data obtained by segmenting the voice data in units of sequences of conversation by use of the speaker models derived by the speaker model deriving means.