Contextual Entity Resolution for Voice-Activated Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice-activated electronic devices face challenges in accurately interpreting user commands when content is being displayed, as they lack contextual information to resolve anaphoric references and fill slots in intents, leading to poor user experience due to the need for additional interactions.
Innovation Solution
The system uses contextual metadata generated from the displayed content, formatted into domain-specific intents and slots, to assist natural language understanding in resolving user intents by matching entities with displayed content, thereby enhancing the accuracy of command interpretation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If voice-activated electronic devices process user commands without contextual information from displayed content, then device complexity is reduced, but intent resolution accuracy deteriorates
Solution Approach 1:
The system performs preliminary actions by generating contextual metadata from displayed content before processing user commands. The metadata generation component creates structured representations of displayed content (text, images, videos) in advance, which are then used during command interpretation to resolve anaphoric references and fill intent slots, improving accuracy without adding complexity during the critical command processing moment
Solution Approach 2:
The patent introduces contextual metadata as an intermediary between displayed content and natural language processing. This metadata acts as a mediator that bridges the gap between visual content and voice commands, providing the necessary context for accurate intent resolution without requiring direct integration between display and processing systems, thus managing complexity
2Measurement precision
If the system requests additional interactions to clarify user intent, then intent resolution accuracy improves, but productivity deteriorates
Solution Approach 1:
The system performs preliminary analysis of displayed content to generate contextual metadata before user commands are issued. This advance preparation enables the natural language processing component to immediately resolve anaphoric references and fill intent slots using pre-extracted contextual information, eliminating the need for follow-up clarification questions and maintaining high interaction efficiency
Solution Approach 2:
The system implements feedback by continuously monitoring displayed content and updating contextual metadata in real-time. This feedback loop ensures that the most current contextual information is available for command interpretation, enabling accurate intent resolution in a single interaction without requiring additional user inputs
3Measurement precision
If contextual metadata is generated and integrated into natural language processing, then intent resolution accuracy improves, but device complexity increases
Solution Approach 1:
The patent segments the processing system into distinct functional components: a metadata generation component that extracts contextual information from displayed content, and a natural language processing component that utilizes this metadata for intent resolution. This segmentation allows each component to specialize in its specific task, improving anaphora resolution accuracy while managing overall system complexity through modular design
Solution Approach 2:
The contextual metadata structure is designed to be universal and domain-agnostic, capable of representing various types of displayed content (text, images, videos) in a unified format. This multi-functionality allows the same metadata processing pipeline to handle diverse content types without increasing complexity, as the underlying structure remains consistent across different domains
Data Source
AI summary
Methods and systems for resolving entities using multi-modal functionality are described herein. Voice activated electronic devices may, in some embodiments, be capable of displaying content using a display screen. Contextual metadata representing the content rendered by the display screen may describe entities having similar attributes as an identified intent from natural language understanding processing. When natural language understanding processing attempts to resolve one or more declared slots for a particular intent, matching slots from the contextual metadata may be determined, and the matching entities may be placed in an intent selected context file to be included with the natural language understanding's output data. The output data may be provided to a corresponding application for causing one or more actions to be performed.


