Voice Processing Apparatus Emotion-Aware Speaker Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice processing techniques struggle to accurately associate a specific speaker with their voice data when the speaker's emotion changes, leading to lower clustering performance and incorrect speaker association.
Innovation Solution
A voice processing method that detects voice sections, calculates feature amounts, determines emotions, and clusters these amounts based on change vectors to accurately associate speakers with their voice data, even when emotions change, by using a combination of feature extraction and clustering techniques such as principal component analysis and k-nearest neighbor methods.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional clustering methods are used to group utterances by acoustic similarity, then the clustering process is simple to implement, but the clustering performance deteriorates when speakers' emotions change
Solution Approach 1:
The patent applies preliminary action by detecting emotions of speakers before performing clustering. The emotion detection unit identifies emotional states of multiple speakers in advance, and this emotion information is then used to guide the clustering process. This allows the system to adapt clustering parameters based on emotional context, improving speaker association accuracy when emotions change during speech.
Solution Approach 2:
The patent implements parameter changes by dynamically adjusting clustering parameters based on detected emotions. When emotions change, the system modifies clustering parameters such as similarity thresholds or weighting factors to account for emotional variations in acoustic features. This enables the clustering process to remain effective despite emotional fluctuations in speakers' voices.
2Measurement precision
If emotion detection is added to improve speaker association, then speaker identification accuracy improves, but processing time and computational complexity increase
Solution Approach 1:
The patent applies segmentation by dividing the voice data processing into distinct stages: emotion detection, feature extraction, and clustering. Each stage processes specific aspects of the data independently, allowing for optimized computation at each step. This segmented approach reduces overall processing time compared to a monolithic processing system while maintaining high speaker identification accuracy through emotion-aware clustering.
Data Source
AI summary
A non-transitory computer-readable recording medium having stored therein a program that causes a computer to execute a procedure, the procedure includes detecting a plurality of voice sections from an input sound that includes voices of a plurality of speakers, calculating a feature amount of each of the plurality of voice sections, determining a plurality of emotions, corresponding to the plurality of voice sections respectively, of a speaker of the plurality of speakers for each of the plurality of voice sections, and clustering a plurality of feature amounts, based on a change vector from the feature amount of the voice section determined as a first emotion of the plurality of emotions of the speaker to the feature amount of the voice section determined as a second emotion of the plurality of emotions different from the first emotion.


