Intent Re-ranker Using Displayed Content Entity Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice-activated electronic devices face challenges in accurately determining the intent of an utterance, especially when multiple intents have similar confidence scores, and lack of contextual information from displayed content hinders precise action execution.
Innovation Solution
The system performs intent re-ranking by incorporating entity data associated with displayed content, using a multi-domain architecture that extracts features from the content to update intent hypothesis scores, and an orchestrator facilitates domain re-ranking to select the most likely intent based on contextual information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional intent determination methods are used, then the system can process utterances, but accuracy is reduced when multiple intents have similar confidence scores
Solution Approach 1:
The system performs preliminary extraction of entity data from displayed content before intent determination. This advance preparation of contextual information allows the re-ranking process to accurately resolve ambiguities when multiple intents have similar confidence scores, thereby improving both measurement precision and reliability of intent selection.
Solution Approach 2:
The system implements a feedback mechanism where entity data from displayed content is fed back into the intent determination process through re-ranking. This feedback loop refines the initial intent hypotheses by incorporating contextual information, ensuring more accurate and reliable intent selection even when confidence scores are similar.
2Manufacturing precision
If contextual information from displayed content is not incorporated, then the system operates simpler, but precision of action execution is hindered
Solution Approach 1:
The system segments the intent determination process into distinct stages: initial intent hypothesis generation, entity data extraction from displayed content, and re-ranking based on contextual information. This segmentation allows the system to incorporate complex contextual processing only when needed, maintaining action execution precision while managing system complexity through modular architecture.
Solution Approach 2:
The system introduces an intermediary component that extracts entity data from displayed content and uses it as a mediator in the intent re-ranking process. This intermediary layer enables precise action execution by bridging the gap between raw utterances and intent determination, without requiring complete system redesign.
3Measurement precision
If intent re-ranking with entity data is implemented, then ambiguity in intent hypotheses is resolved, but processing complexity increases
Solution Approach 1:
The system performs preliminary extraction of entity data from displayed content before the re-ranking process. This advance preparation organizes contextual information in a readily usable format, enabling efficient discrimination between intent hypotheses without requiring complex processing during the actual re-ranking stage.
Solution Approach 2:
The system changes the ranking parameters of intent hypotheses by incorporating entity data from displayed content. This parameter adjustment allows precise discrimination between ambiguous intents by re-evaluating hypotheses based on contextual relevance, achieving high precision without proportionally increasing processing complexity.
Data Source
AI summary
Methods and systems for determining an intent of an utterance using contextual information associated with a requesting device are described herein. Voice activated electronic devices may, in some embodiments, be capable of displaying content using a display screen. Entity data representing the content rendered by the display screen may describe entities having similar attributes as an identified intent from natural language understanding processing. Natural language understanding processing may attempt to resolve one or more declared slots for a particular intent and may generate an initial list of intent hypotheses ranked to indicate which are most likely to correspond to the utterance. The entity data may be compared with the declared slots for the intent hypotheses, and the list of intent hypothesis may be re-ranked to account for matching slots from the contextual metadata. The top ranked intent hypothesis after re-ranking may then be selected as the utterance's intent.


