Disambiguation Model for Voice-Based Item Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computing devices face challenges in accurately interpreting and selecting items on a screen using voice input due to limitations in rule-based grammars, leading to a time-consuming trial-and-error approach for users.
Innovation Solution
A model-based approach is implemented, where a disambiguation model is applied to utterances to identify corresponding items on the screen by extracting referential features, utilizing statistical classifiers, semantic parsers, and location parsers to determine the correct item for user queries.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If rule-based grammars are used for voice input interpretation, then the system structure is simple, but the accuracy of item identification is poor and user experience deteriorates
Solution Approach 1:
The patent transforms the voice input system from rule-based grammar matching to a model-based approach using statistical classifiers and semantic parsers. This parameter change in the interpretation methodology enables the system to handle diverse natural language queries accurately while extracting referential features for item identification, resolving the contradiction between accuracy and complexity.
Solution Approach 2:
The patent replaces the mechanical rule-based grammar system with a statistical model-based system. Instead of relying on predefined grammatical rules, the system uses statistical classifiers and semantic parsers to interpret voice inputs, achieving higher accuracy in item identification while managing complexity through model-based processing.
2Adaptability or versatility
If rule-based grammars with limited commands are used, then the system complexity is low, but the adaptability to different user queries is poor
Solution Approach 1:
The patent implements a universal disambiguation model that can handle multiple types of natural language queries through a single system architecture. The semantic parser and referential feature extraction mechanisms enable the system to adapt to various query formats and expressions, providing versatile item identification without requiring separate rule sets for different command types.
Solution Approach 2:
The patent introduces dynamic adaptability through the disambiguation model that can process and interpret diverse natural language expressions. The system dynamically extracts referential features from different types of queries and adapts its interpretation based on the extracted features, enabling versatile handling of user inputs while maintaining a unified system structure.
3Productivity
If trial-and-error approach is used by users, then the system requirements are simple, but the time consumption increases
Solution Approach 1:
The patent implements a feedback mechanism through the disambiguation model that analyzes user voice inputs and provides accurate item identification in real-time. The system uses extracted referential features to disambiguate queries and immediately identify the intended item, providing feedback that eliminates the need for trial-and-error attempts and significantly reduces selection time.
Solution Approach 2:
The patent performs preliminary action by pre-processing and extracting referential features from voice inputs before item selection. The disambiguation model prepares the interpretation in advance by analyzing the query structure and extracting key features, enabling rapid and accurate item identification without requiring multiple user attempts.
Data Source
AI summary
A model-based approach for on-screen item selection and disambiguation is provided. An utterance may be received by a computing device in response to a display of a list of items for selection on a display screen. A disambiguation model may then be applied to the utterance. The disambiguation model may be utilized to determine whether the utterance is directed to at least one of the list of displayed items, extract referential features from the utterance and identify an item from the list corresponding to the utterance, based on the extracted referential features. The computing device may then perform an action which includes selecting the identified item associated with utterance.


