Actionable Suggestion Interface for Visual Entity Disambiguation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for detecting and disambiguating named entities in visual inputs, such as images, are limited by the lack of textual context, leading to difficulties in identifying entities like locations, organizations, and actions, and failing to provide actionable suggestions.
Innovation Solution
The proposed solution involves using a device and method that detects user selections of textual elements and correlates them with non-textual elements using a layout learning model, to derive entities and provide actionable suggestions through a user-operable interface.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If only textual content is extracted from images for entity detection, then the process is simple, but entity detection accuracy and disambiguation capability deteriorate due to lack of contextual information
Solution Approach 1:
The patent merges textual content extraction with visual element detection and correlation. The system extracts text from images, detects visual elements, and correlates them based on spatial relationships and layout patterns to derive entities. This combination resolves the contradiction by maintaining the simplicity of text extraction while adding visual context to improve entity detection accuracy and disambiguation capability.
Solution Approach 2:
The patent introduces a layout learning model as an intermediary that connects textual elements and visual elements. This intermediary component correlates text with visual elements based on their spatial relationships and layout patterns, enabling accurate entity detection without requiring complex direct processing. The layout learning model acts as a mediator that bridges the gap between simple text extraction and accurate entity recognition.
2Ease of manufacture
If predefined entity categories are used for named entity recognition, then the classification process is straightforward, but the system cannot detect new or ambiguous entity types like locations and organizations in visual context
Solution Approach 1:
The patent applies dynamics by making the entity detection process adaptive rather than static. The layout learning model dynamically correlates textual elements with visual elements based on their spatial relationships and layout patterns. This dynamic correlation enables the system to detect new and ambiguous entity types such as locations and organizations in visual context, while maintaining the simplicity of classification through learned patterns rather than rigid predefined categories.
3Measurement precision
If long portions of text are used for entity detection, then detection accuracy improves, but the system becomes less effective for scene text and small text scenarios with limited contextual information
Solution Approach 1:
The patent applies segmentation by dividing the text and visual content into discrete elements and their spatial relationships. Instead of processing long portions of text as a whole, the system segments the content into individual textual elements and visual elements, then correlates them based on their spatial relationships. This segmentation enables accurate entity detection in scene text and small text scenarios where limited contextual information is available, as the system can identify entities through local spatial correlations rather than requiring extensive text context.
Data Source
AI summary
An electronic device and method performed by the electronic device for suggesting at least one are provided. The method includes detecting a user selection of a first modality in a display of the electronic device and at least one second modality present in vicinity to the first modality, deriving at least one pair of the first modality and at least one second modality by correlating the first modality and the at least one second modality, and providing a suggestion in a form of a user-operable interface based on the derived at least one pair. A user operation on the suggestion initiates execution of the at least one action on the first modality via an application.


