Urgency Estimation via Speech Speed and Pitch Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional call urgency level estimation techniques are limited as they rely on specific words, such as 'Help', and fail to estimate urgency from speech that does not include these words.
Innovation Solution
An urgency level estimation apparatus that extracts features like speaking speed, voice pitch, and power level from uttered speech, using statistical values and machine learning models to determine the urgency level without requiring specific words.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If vocal tract feature amount of specific word is used for urgency estimation, then estimation accuracy for words like 'Help' is improved, but applicability to speech not including specific words deteriorates
Solution Approach 1:
The patent changes the parameters used for urgency estimation from specific word-based vocal tract features to general speech features including speaking speed, voice pitch, and power level. This allows the system to estimate urgency from any speech content rather than requiring specific keywords, thereby improving adaptability while maintaining estimation accuracy through multiple complementary features.
2Adaptability or versatility
If multiple feature amounts (speaking speed, voice pitch, power level) are extracted and analyzed, then adaptability to various speech types is improved, but device complexity increases
Solution Approach 1:
The patent segments the urgency estimation task into multiple independent feature extraction components: speaking speed extraction, voice pitch extraction, and power level extraction. Each component processes a specific aspect of the speech signal independently, then their results are integrated for final urgency determination. This segmentation approach manages complexity by breaking down the overall system into manageable, specialized modules.
Data Source
AI summary
An urgency level estimation technique of estimating an urgency level of a speaker for free uttered speech, which does not require a specific word, is provided. An urgency level estimation apparatus includes a feature amount extracting part configured to extract a feature amount of an utterance from uttered speech, and an urgency level estimating part configured to estimate an urgency level of a speaker of the uttered speech from the feature amount based on a relationship between a feature amount extracted from uttered speech and an urgency level of a speaker of the uttered speech, the relationship being determined in advance, and the feature amount includes at least one of a feature indicating speaking speed of the uttered speech, a feature indicating voice pitch of the uttered speech and a feature indicating a power level of the uttered speech.


