Voice Command Intent Prediction Using Stored Context
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition systems in electronic devices struggle to accurately interpret user intentions from imperfect or free-style voice commands, often failing to execute the intended operation or performing a different action than intended.
Innovation Solution
An electronic apparatus that accumulates situation information based on user interactions and uses this data to predict user intentions by analyzing factors such as device, space, time, and utterance similarity, allowing it to identify and execute the correct operation even with incomplete voice commands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the electronic apparatus uses traditional voice recognition methods to interpret user commands, then the system structure remains simple, but the accuracy of interpreting user intentions from imperfect or free-style voice commands deteriorates
Solution Approach 1:
The system performs preliminary actions by storing situation information (device states, space information, time information) before voice recognition occurs. This pre-stored context enables the system to accurately interpret imperfect voice commands by comparing them against known situations, resolving the contradiction between recognition accuracy and system complexity.
Solution Approach 2:
The patent introduces situation information as an intermediary element between the user's voice input and the system's command interpretation. This intermediary layer contains pre-stored contextual data that mediates the matching process, allowing accurate interpretation of free-style commands without significantly increasing system complexity.
2Reliability
If the electronic apparatus stores and processes situation information to predict user intentions, then the accuracy of command execution improves, but the loss of information processing time increases
Solution Approach 1:
The system performs preliminary action by pre-storing situation information including device states, space information, and time information before recognition occurs. This advance preparation enables fast matching during recognition without requiring complex real-time analysis, thus improving reliability while minimizing processing time loss.
Solution Approach 2:
The patent creates copies of situation information in a structured format that can be quickly matched against voice commands. By maintaining pre-processed situation data as references, the system avoids time-consuming analysis during actual command recognition, balancing reliability with processing speed.
3Measurement precision
If the electronic apparatus requires complete information in voice commands for accurate operation, then the precision of command interpretation improves, but the ease of operation deteriorates as users must speak formally structured sentences
Solution Approach 1:
The patent introduces situation information as an intermediary that bridges the gap between simple user speech and precise command interpretation. The pre-stored situation data (device contexts, space information, time information) acts as a mediator that completes the meaning of imperfect commands, allowing users to speak naturally while maintaining high interpretation precision.
Solution Approach 2:
The system performs self-service by automatically retrieving and applying relevant situation information from its stored database without requiring users to provide complete command structures. The electronic apparatus independently fills in missing information from its situation database, enabling users to speak casually while the system handles the complexity of complete command interpretation.
Data Source
AI summary
An electronic apparatus includes a communicator configured to communicate with a plurality of external apparatus. A storage is configured to store situation information. A processor is configured to, based on a first utterance of a user, control a first operation corresponding to the first utterance to be carried out from among a plurality of operations related to the plurality of external apparatuses. Situation information corresponding to each of a plurality of situations where the first operation is carried out based on the first utterance is stored in the storage. Based on a second utterance of the user, a second operation is identified corresponding to the second utterance from among the plurality of operations based on the stored situation information, and the identified second operation is carried out.


