Utterance State Detection via Voice Fluctuation Statistics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing emotion detection techniques require preparation of user-specific reference information, making them cumbersome and limiting to a specific user, and thus not applicable without prior setup.

Innovation Solution

An utterance state detection device that extracts high-frequency elements from user voice stream data, calculates fluctuation degrees, and uses statistical analysis to detect the utterance state without pre-prepared reference information for each user, utilizing a frequency-analyzing unit, fluctuation degree calculation unit, and statistic calculation unit.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If user-specific reference information is prepared in advance for each user, then emotion detection accuracy is improved, but device complexity and preparation time increase

Engineering Contradiction:
Improveemotion detection accuracyVSAvoidreference information preparation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs self-calibration by automatically analyzing the user's voice characteristics during initial use and generating personalized reference information without requiring manual setup. The device serves itself by autonomously creating the user-specific parameters needed for accurate emotion detection, eliminating the need for external preparation work while maintaining high detection accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary voice analysis and reference information generation automatically during the first interaction with each user. By conducting this preparation action in advance during normal operation rather than requiring separate setup procedures, the system ensures accurate emotion detection is ready when needed without adding complexity to the user-facing interface.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If user-specific reference information is prepared in advance for each user, then emotion detection accuracy is improved, but ease of operation deteriorates

Engineering Contradiction:
Improveemotion detection accuracyVSAvoidease of technique application
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system automatically performs user-specific calibration without requiring user intervention or technical setup procedures. Users simply begin using the device, and the system autonomously adapts to their voice characteristics, making the technique as easy to apply as any other voice analysis tool while maintaining personalized accuracy.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If statistics are calculated from a large number of fluctuation degrees, then detection accuracy is improved, but processing time increases

Engineering Contradiction:
Improveutterance state detection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system pre-calculates and stores statistical parameters from voice fluctuation data during periods when processing requirements are lower. By performing this computationally intensive analysis in advance and storing the results, the system can quickly retrieve pre-computed statistics during actual emotion detection without real-time processing delays, achieving both accuracy and speed.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9099088B2Utterance state detection device and utterance state detection method
Publication Date: 2015.08.04 FUJITSU LTD
  • US9099088B2 patent drawing
  • US9099088B2 patent drawing
  • US9099088B2 patent drawing

AI summary

An utterance state detection device includes an user voice stream data input unit that gets user voice stream data of an user, a frequency element extraction unit that extracts high frequency elements by frequency-analyzing the user voice stream data, a fluctuation degree calculation unit that calculates a fluctuation degree of the high frequency elements thus extracted every unit time, a statistic calculation unit that calculates a statistic every certain interval based on a plurality of the fluctuation degrees in a certain period of time, and an utterance state detection unit that detects an utterance state of a specified user based on the statistic obtained from user voice stream data of the specified user.