Multi-Intent Voice Processing via Segmentation and Preliminary Action

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing chat-bots and voice assistants often fail to understand user queries that consist of multiple sentences or intents, providing inaccurate or inappropriate responses due to their inability to process verbose user input effectively.

Innovation Solution

A terminal device and server system that transmits and processes user voice inputs with multiple intents, using word use information and user-related information to identify response orders and provide accurate, sequential responses based on importance and relevance, with the option to modify responses based on user feedback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a chat-bot determines one speech as one intent and provides a predetermined single answer, then the system complexity is low and operation is simple, but the accuracy of response is insufficient when user input contains multiple intents

Engineering Contradiction:
Improveresponse accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the user's speech into multiple intents by analyzing word use information and user-related information. The server divides the speech processing task into identifying individual intents, determining their response orders, and generating separate responses for each intent, thereby achieving accurate handling of multi-intent queries without overwhelming system complexity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary action by pre-identifying multiple intents and their response orders before generating the final response. The server determines the sequence in which to address each intent based on word use information and user-related information, preparing a structured response plan in advance that improves overall response accuracy

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If a chat-bot processes verbose user input with multiple sentences, then the response accuracy improves, but the processing time and complexity increase

Engineering Contradiction:
Improveresponse accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments verbose user input into discrete intents, allowing parallel processing of multiple meaning units. By identifying and separating different intents within the same speech, the system can process them efficiently according to determined response orders, reducing overall processing time while maintaining accuracy

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary identification of multiple intents and their response orders before full response generation. This advance structuring of the response plan based on word use information and user-related information enables more efficient processing of verbose input, reducing the time required for complete response generation

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11538476B2Terminal device, server and controlling method thereof
Publication Date: 2022.12.27 SAMSUNG ELECTRONICS CO LTD
  • US11538476B2 patent drawing
  • US11538476B2 patent drawing
  • US11538476B2 patent drawing

AI summary

A terminal device is provided and includes a communication interface including circuitry, a display and at least one processor configured to control the communication interface to transmit a user voice including a plurality of intents to an external server, based on word use information included in the user voice and summary information regarding the user voice generated based on user-related information being received from the external server, control the display to display the received summary information, based on a user feedback regarding the summary information being input, transmit information regarding the user feedback to the external server, and based on response information regarding the user voice generated based on the user feedback being received from the external server, control the display to provide the response information.