Utterance State Detection via Voice Fluctuation Statistics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing emotion detection techniques require preparation of user-specific reference information, making them cumbersome and limiting to a specific user, and thus not applicable without prior setup.
Innovation Solution
An utterance state detection device that extracts high-frequency elements from user voice stream data, calculates fluctuation degrees, and uses statistical analysis to detect the utterance state without pre-prepared reference information for each user, utilizing a frequency-analyzing unit, fluctuation degree calculation unit, and statistic calculation unit.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If user-specific reference information is prepared in advance for each user, then emotion detection accuracy is improved, but device complexity and preparation time increase
Solution Approach 1:
The system performs self-calibration by automatically analyzing the user's voice characteristics during initial use and generating personalized reference information without requiring manual setup. The device serves itself by autonomously creating the user-specific parameters needed for accurate emotion detection, eliminating the need for external preparation work while maintaining high detection accuracy.
Solution Approach 2:
The system performs preliminary voice analysis and reference information generation automatically during the first interaction with each user. By conducting this preparation action in advance during normal operation rather than requiring separate setup procedures, the system ensures accurate emotion detection is ready when needed without adding complexity to the user-facing interface.
2Measurement precision
If user-specific reference information is prepared in advance for each user, then emotion detection accuracy is improved, but ease of operation deteriorates
Solution Approach 1:
The system automatically performs user-specific calibration without requiring user intervention or technical setup procedures. Users simply begin using the device, and the system autonomously adapts to their voice characteristics, making the technique as easy to apply as any other voice analysis tool while maintaining personalized accuracy.
3Measurement precision
If statistics are calculated from a large number of fluctuation degrees, then detection accuracy is improved, but processing time increases
Solution Approach 1:
The system pre-calculates and stores statistical parameters from voice fluctuation data during periods when processing requirements are lower. By performing this computationally intensive analysis in advance and storing the results, the system can quickly retrieve pre-computed statistics during actual emotion detection without real-time processing delays, achieving both accuracy and speed.
Data Source
AI summary
An utterance state detection device includes an user voice stream data input unit that gets user voice stream data of an user, a frequency element extraction unit that extracts high frequency elements by frequency-analyzing the user voice stream data, a fluctuation degree calculation unit that calculates a fluctuation degree of the high frequency elements thus extracted every unit time, a statistic calculation unit that calculates a statistic every certain interval based on a plurality of the fluctuation degrees in a certain period of time, and an utterance state detection unit that detects an utterance state of a specified user based on the statistic obtained from user voice stream data of the specified user.


