Voice System Utterance Segmentation for Partial Repetition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice recognition systems fail to accurately respond to user requests for reuttering specific parts of system utterances, leading to unnecessary repetition and incorrect information being provided when users ask for clarification or specific details.
Innovation Solution
An information processing device and method that determine the type of user utterance, such as requesting all or part of the previous system utterance, to generate appropriate responses, including reuttering all or a portion of the previous system utterance based on the user's request.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the system reutters all system utterances from the beginning when a user requests reutterance, then the user can hear the complete information again, but the user has to repeatedly hear unnecessary information, resulting in wasted time
Solution Approach 1:
The patent segments the system utterance into multiple parts or sections, allowing the system to identify and reutter only the specific portion that the user wants to hear again, rather than reuttering the entire utterance from the beginning. This segmentation enables precise control over what information is reuttered, eliminating unnecessary repetition while maintaining completeness of relevant information.
Solution Approach 2:
The patent extracts or identifies the specific portion of the system utterance that the user wants to reutter by analyzing the user's request. The system takes out only the relevant segment from the complete system utterance and reutters that extracted portion, avoiding the waste of time from hearing unnecessary information while ensuring the complete relevant information is provided.
2Ease of manufacture
If the system uses predefined keywords to detect reutterance requests, then the system can process specific user utterances, but the system cannot respond to cases where these specific words are not included
Solution Approach 1:
The patent changes the detection parameter from relying on specific predefined keywords to analyzing the temporal relationship and semantic content between the user utterance and the system utterance. By changing how the system detects reutterance requests - focusing on the relationship between utterances rather than specific words - the system gains versatility to respond to various user utterances while maintaining a relatively simple detection mechanism based on timing and content analysis.
3Adaptability or versatility
If the system answers questions using a general document set, then the system can provide information, but the system may output an answer different from what the user wanted to hear
Solution Approach 1:
The patent applies local quality by focusing the information retrieval on the specific local context of the system utterance that the user wants to reutter, rather than searching through a general document set. The system retrieves information locally from the recent system utterance with matching semantic content, ensuring high accuracy in matching user intent while maintaining the ability to answer various questions about the system's own output.
Data Source
AI summary
A device and a method that determine an utterance type of a user utterance and generate a system response according to a determination result are achieved. A user utterance type determination unit that determines an utterance type of a user utterance, and a system response generation unit that generates a system response according to a type determination result determined by the user utterance type determination unit are included. The user utterance type determination unit determines whether the user utterance is of type A that requests all reutterances of a system utterance immediately before the user utterance, or type B that requests a reutterance of a part of the system utterance immediately before the user utterance. The system response generation unit generates a system response to reutter all system utterances immediately before the user utterance in a case where the user utterance is of type A, and generates a system response to reutter a part of system utterances immediately before the user utterance in a case where the user utterance is of type B.


