Voice Assistant Input Processing via Instruction Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition technologies are limited in performing a wide range of user control operations and fail to handle complex user inputs effectively, requiring a method to process and distribute user inputs across multiple voice assistants efficiently.
Innovation Solution
A method that converts user voice inputs into instructions, divides complex instructions into partial ones, determines their domains, and distributes them to the most suitable voice assistants based on reliability, performance, and frequency of use, allowing for sequential or parallel processing as needed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech recognition technology is used to recognize user speech, then the device can perform control operations, but the range of functions is limited and cannot handle complex user inputs effectively
Solution Approach 1:
The patent segments a complex user instruction into multiple partial instructions, each corresponding to a specific domain or task. For example, a complex instruction like 'set the temperature to 25 degrees and turn on the lights' is divided into separate sub-instructions for temperature control and lighting control. This segmentation allows each partial instruction to be processed by specialized voice assistants with specific expertise, thereby expanding the range of functions while maintaining reliability in handling complex inputs.
2Adaptability or versatility
If a single voice assistant processes all user inputs, then the system structure is simple, but the capability to perform varied tasks is limited
Solution Approach 1:
The patent creates a universal instruction processing framework that can handle diverse tasks through multiple specialized voice assistants. The main processing unit acts as a universal coordinator that receives any type of user input, analyzes it, and routes it to the appropriate specialized assistant. This multi-functionality approach enables the system to perform varied tasks (weather queries, device control, information search, etc.) while maintaining a relatively simple overall structure through standardized processing protocols.
3Productivity
If complex instructions are processed as a whole, then the processing logic is straightforward, but the efficiency and accuracy of task execution decreases
Solution Approach 1:
The patent applies segmentation by breaking down complex instructions into smaller, manageable partial instructions. Each partial instruction is then processed independently by appropriate voice assistants, which improves execution efficiency and accuracy. The processing logic becomes more complex at the decomposition stage but simpler at the execution stage, as each assistant handles only its specific domain with straightforward, optimized logic.
4Reliability
If voice assistants are selected based on multiple criteria (reliability, performance, frequency), then the task execution accuracy improves, but the selection process becomes more complex
Solution Approach 1:
The patent implements a feedback mechanism where the system continuously monitors and evaluates the performance of different voice assistants based on multiple criteria including reliability, performance metrics, and usage frequency. This feedback information is used to dynamically adjust the selection and weighting of assistants for different tasks. The feedback loop ensures high execution accuracy while the selection process complexity is managed through automated evaluation algorithms that learn from historical performance data.
Data Source
AI summary
Provided is a method of processing a user input to deliver the user input to at least one of a plurality of assistants, includes: converting a user input including a voice signal based on a predetermined rule to generate an instruction; splitting a complex instruction into partial instructions based on that the generated instruction is the complex instruction requesting two or more events; and determining a domain of each of the partial instructions and distributing the partial instructions to at least one of a plurality of voice assistants based on the domain. According to an embodiment, the washer may be related to artificial intelligence (AI) modules, unmanned aerial vehicles (UAVs), robots, augmented reality (AR) devices, virtual reality (VR) devices, and 5G service-related devices.


