Multi-Intent Voice Processing via Segmentation and Preliminary Action
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing chat-bots and voice assistants often fail to understand user queries that consist of multiple sentences or intents, providing inaccurate or inappropriate responses due to their inability to process verbose user input effectively.
Innovation Solution
A terminal device and server system that transmits and processes user voice inputs with multiple intents, using word use information and user-related information to identify response orders and provide accurate, sequential responses based on importance and relevance, with the option to modify responses based on user feedback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a chat-bot determines one speech as one intent and provides a predetermined single answer, then the system complexity is low and operation is simple, but the accuracy of response is insufficient when user input contains multiple intents
Solution Approach 1:
The patent segments the user's speech into multiple intents by analyzing word use information and user-related information. The server divides the speech processing task into identifying individual intents, determining their response orders, and generating separate responses for each intent, thereby achieving accurate handling of multi-intent queries without overwhelming system complexity
Solution Approach 2:
The patent performs preliminary action by pre-identifying multiple intents and their response orders before generating the final response. The server determines the sequence in which to address each intent based on word use information and user-related information, preparing a structured response plan in advance that improves overall response accuracy
2Measurement precision
If a chat-bot processes verbose user input with multiple sentences, then the response accuracy improves, but the processing time and complexity increase
Solution Approach 1:
The patent segments verbose user input into discrete intents, allowing parallel processing of multiple meaning units. By identifying and separating different intents within the same speech, the system can process them efficiently according to determined response orders, reducing overall processing time while maintaining accuracy
Solution Approach 2:
The patent performs preliminary identification of multiple intents and their response orders before full response generation. This advance structuring of the response plan based on word use information and user-related information enables more efficient processing of verbose input, reducing the time required for complete response generation
Data Source
AI summary
A terminal device is provided and includes a communication interface including circuitry, a display and at least one processor configured to control the communication interface to transmit a user voice including a plurality of intents to an external server, based on word use information included in the user voice and summary information regarding the user voice generated based on user-related information being received from the external server, control the display to display the received summary information, based on a user feedback regarding the summary information being input, transmit information regarding the user feedback to the external server, and based on response information regarding the user voice generated based on the user feedback being received from the external server, control the display to provide the response information.


