Speaker Co-occurrence Model for Multi-Speaker Voice Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice data analyzing devices fail to consider the relationships among speakers, leading to decreased recognition accuracy, particularly in scenarios involving multiple speakers, such as criminal investigations or phishing scams.
Innovation Solution
A voice data analyzing device that derives speaker models and co-occurrence models from segmented voice data, incorporating the relationships between speakers through speaker model learning and co-occurrence model derivation, enabling accurate recognition of speakers in conversations involving multiple individuals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speaker models are learned independently for each speaker using voice data and speaker labels, then the speaker recognition process can be executed for each speaker model independently, but the relationship between speakers cannot be utilized leading to deteriorated recognition accuracy
Solution Approach 1:
The patent merges independent speaker model learning with relationship analysis by introducing a speaker relationship model that captures co-occurrence patterns between speakers. The system combines speaker-specific acoustic models with a relationship layer that models how speakers interact and co-appear in conversations, thereby utilizing relationship information to improve recognition accuracy while maintaining the modular structure of independent speaker model learning.
Solution Approach 2:
The patent adds another dimension to the traditional speaker recognition framework by incorporating speaker relationship information as a separate layer. Instead of only modeling speakers independently in the acoustic feature space, the system extends the representation to include relationship features that capture co-occurrence patterns, effectively adding a relational dimension to the recognition problem.
2Adaptability or versatility
If clustering of learned speakers is conducted by determining vocal tract length expansion/contraction coefficient, then speakers can be grouped based on physical characteristics, but the relationship between speakers is still not discussed leading to limited recognition capability
Solution Approach 1:
The patent changes the parameters used for speaker representation by introducing relationship-based features alongside traditional acoustic parameters. Instead of relying solely on vocal tract length coefficients for clustering, the system incorporates speaker relationship models that capture co-occurrence patterns, thereby changing the parameter space to include relational information that improves recognition capability.
Data Source
AI summary
A voice data analyzing device comprises speaker model deriving means which derives speaker models as models each specifying character of voice of each speaker from voice data including a plurality of utterances to each of which a speaker label as information for identifying a speaker has been assigned and speaker co-occurrence model deriving means which derives a speaker co-occurrence model as a model representing the strength of co-occurrence relationship among the speakers from session data obtained by segmenting the voice data in units of sequences of conversation by use of the speaker models derived by the speaker model deriving means.


