Voice Interaction Topic Change and Prosody Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice interaction systems face challenges in detecting 'ask-again' requests, particularly for voices that do not include interjections, leading to processing delays and limited detection capabilities.
Innovation Solution
A voice interaction system that utilizes topic detection and prosodic analysis to identify 'ask-again' requests in user voices, without requiring pre-registration of specific words, by estimating topic changes and analyzing prosodic information to detect significant changes in tone, thereby enabling detection across a wider range of voices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If voice recognition is performed using registered words only, then recognition accuracy is improved, but detection capability is limited to registered words and processing time increases
Solution Approach 1:
The system changes the detection parameters from word-based recognition to prosody-based detection. By analyzing acoustic features such as pitch, energy, and duration of utterances, the system can detect ask-again intent without being constrained by registered vocabulary, thus improving both detection capability and reducing processing time
Solution Approach 2:
The patent replaces the mechanical word-matching system with an acoustic analysis system. Instead of comparing recognized words against a registered dictionary, the system substitutes this with analysis of prosodic features (pitch contours, energy variations, duration patterns) to detect ask-again intent, enabling detection of unregistered words
2Ease of operation
If interjections are used to detect ask-again, then detection simplicity is improved, but detection scope is limited to voices containing interjections
Solution Approach 1:
The system creates a universal detection mechanism that works for all types of ask-again expressions, not just those containing interjections. By analyzing prosodic patterns that are common across different speech acts, the system achieves multi-functionality in detecting various forms of ask-again requests including questions, repetitions, and clarifications
Solution Approach 2:
The system changes from detecting specific linguistic content (interjections) to detecting acoustic pattern parameters (prosody). This parameter shift allows the system to identify ask-again intent through pitch, energy, and duration characteristics that appear across diverse speech types, greatly expanding detection scope
3Measurement precision
If topic change detection is combined with prosodic analysis, then ask-again detection accuracy is improved, but system complexity increases
Solution Approach 1:
The system segments the ask-again detection task into two independent modules: topic change detection and prosodic analysis. Each module processes specific features independently, and their results are combined to make the final detection decision. This segmentation improves accuracy while managing complexity through modular design
Solution Approach 2:
The system introduces topic change detection as an intermediary condition that triggers or modulates prosodic analysis. When a topic change is detected, the system pays closer attention to prosodic features, using the topic information as a mediator to enhance the sensitivity and accuracy of ask-again detection without continuously processing all prosodic data
Data Source
AI summary
A voice interaction system performs a voice interaction with a user. The voice interaction system includes: topic detection means for estimating a topic of the voice interaction and detecting a change in the topic that has been estimated; and ask-again detection means for detecting, when the change in the topic has been detected by the topic detection means, the user's voice as ask-again by the user based on prosodic information on the user's voice.


