Speaker Selection via Acoustic Feature Likelihood Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speaker adaptive model creating devices fail to adapt to temporal changes in conversations or calls, leading to either comprehensive models that are unstable or local models that are inferior in stability, as they either create a single adaptive model for all utterances or select speakers only once.
Innovation Solution
A speaker selecting device that calculates long-time and short-time likelihoods of speaker models based on voice signals, using a first and second selection mechanism to narrow down speaker selections, and creates adaptive models by integrating sufficient statistics from both, allowing for accurate and stable speaker selection per utterance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a single comprehensive adaptive model is created for all utterances, then the model covers all speakers, but the model becomes unstable due to temporal changes in acoustic features
Solution Approach 1:
The patent divides the single comprehensive model into multiple local adaptive models, each corresponding to a specific speaker. Instead of creating one model to cover all speakers, the system creates separate models for each speaker identified in the utterance, allowing each model to be optimized for its specific speaker's acoustic characteristics while adapting to temporal changes.
Solution Approach 2:
The patent introduces dynamic speaker selection where the set of adaptive models is not fixed but changes based on the detected speaker in each utterance. The system dynamically determines which speakers to include in the adaptive model creation process based on real-time acoustic feature analysis, allowing the model to adapt to temporal variations in speaker characteristics.
2Device complexity
If speakers are selected only once for the entire conversation, then the selection process is simple, but the selection cannot adapt to temporal changes in acoustic features
Solution Approach 1:
The patent implements periodic speaker selection where speakers are re-evaluated and re-selected at regular intervals or for each utterance. Instead of a one-time selection, the system periodically analyzes acoustic features and updates the speaker identification, allowing the selection to adapt to temporal changes while maintaining a manageable selection process.
Solution Approach 2:
The patent introduces feedback mechanisms where the acoustic features extracted from utterances are used to continuously refine and update the speaker identification and model selection. The system uses the detected acoustic characteristics to feedback into the speaker selection process, enabling adaptive selection that responds to temporal variations in speaker acoustic features.
3Adaptability or versatility
If local adaptive models are created for each speaker, then temporal changes are adapted to, but the model stability decreases due to insufficient data for each local model
Solution Approach 1:
The patent merges the advantages of both comprehensive and local models by creating a hybrid approach where speaker-specific adaptive models are combined with a comprehensive speaker list. The system combines the detailed acoustic modeling for each speaker with the broader context of all speakers in the conversation, ensuring sufficient data for each local model while maintaining overall reliability through the comprehensive framework.
Solution Approach 2:
The patent applies preliminary actions by pre-identifying and pre-modeling speakers before the main conversation processing. The system performs preliminary speaker detection and model creation using initial utterances, building up sufficient statistics and acoustic models in advance, which then provide a stable foundation for subsequent local adaptive models during the conversation.
Data Source
AI summary
To enable selection of a speaker, the acoustic feature value of which is similar to that of an utterance speaker, with accuracy and stability, while adapting to changes even when the acoustic feature value of the speaker changes every moment, a long-time speaker score is calculated (log likelihood of each of a plurality of speaker models stored in a speaker model storage with respect to the acoustic feature value) based on an arbitrary number of utterances, for example, and a short-time speaker score is calculated based on a short-time utterance, for example. Speakers are selected corresponding to a predetermined number of speaker models having a high long-time speaker score. Speakers are selected corresponding to the speaker models, the number of which is smaller than the predetermined number and the short-time speaker sore of which is high, from among the speakers having a high long-time speaker score.


