Intelligent Speech Processing System for Multi-Intent Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition services fail to accurately process user inputs, especially when insufficient information is provided, leading to difficulties in grasping the user's intent and failing to execute multiple application requests, as they only display results corresponding to voice inputs without processing the intent behind the utterance.
Innovation Solution
An electronic apparatus and method that uses a rule-based system to process user utterances by arranging states corresponding to the user input, involving a processor, memory, and communication circuitry to transmit data to an external server, receive responses, and display a user interface for further input, allowing for the execution of tasks and filling text input boxes with parameters, thereby enhancing the recognition of user intents and providing a more accurate service.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional speech recognition service displays only the result corresponding to voice input, then the service is simple to operate, but it fails to process user inputs requiring multiple application executions or sufficient information provision
Solution Approach 1:
The patent segments the speech recognition process into distinct stages: initial voice input reception, result display with input method selection options, and follow-up input reception. This segmentation allows the system to handle complex multi-application requests by breaking them down into manageable interaction steps, where users can select different input methods (voice, text, keyboard) at different stages of the task execution流程
Solution Approach 2:
The system dynamically adapts its behavior based on the type of user input received. When a user requests multiple application executions or provides insufficient information, the system transitions from a simple result-display mode to an interactive mode that presents multiple input method options, allowing the service complexity to adjust according to task requirements
2Measurement precision
If conventional speech recognition service processes only voice input, then the input method is simple, but it fails when user voice does not provide sufficient information
Solution Approach 1:
The patent implements a universal input handling framework that supports multiple input methods (voice recognition, text input, keyboard input) within a single speech recognition service. This multi-functionality allows the system to accurately recognize user intent regardless of whether the initial voice input was sufficient, by allowing users to supplement or correct input through alternative methods
Solution Approach 2:
The system provides feedback to users when voice input is insufficient for accurate intent recognition. After displaying the initial recognition result, the system offers feedback options that allow users to select alternative input methods or provide additional information, creating a feedback loop that improves recognition accuracy through iterative input refinement
3Productivity
If conventional speech recognition service executes only single tasks, then the task execution is straightforward, but it cannot handle requests to execute multiple applications
Solution Approach 1:
The patent implements preliminary action by allowing users to declare multiple task execution requests in advance through a single voice input or text input. The system parses and stores these multiple task requests, then executes them in the predetermined sequence without requiring separate interaction steps for each task, thereby improving productivity while managing complexity through upfront task specification
4Reliability
If conventional speech recognition service provides only basic result display, then the interface is simple, but it cannot provide service that matches complex user intent
Solution Approach 1:
The patent introduces an intermediary processing layer between voice input and result execution. This intermediary component analyzes the user input, determines the appropriate input method, processes the input through the selected method, and coordinates multiple tasks if needed. This intermediary structure improves service reliability by ensuring accurate intent recognition while managing processing complexity through a dedicated coordination layer
Data Source
AI summary
A system, apparatus, and method for generating one or more path rules or combination or path rules for performing the operations of an application. According to various embodiments, a user terminal through an intelligence server may generate a path rule and may perform the operation of an application based on the path rule to provide a service. The user terminal may provide a user with the processing status of a user utterance and may receive an additional user input to process the user utterance. For example, when a keyword (or parameter) required to process the user utterance is insufficient, the user terminal may provide the user with a state requiring additional information and may output feedback corresponding to the insufficient state in order to receive the necessary information from the user.


