Entity Fingerprinting via Graph Co-occurrence
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing entity-centric models primarily rely on structured content and fail to effectively represent entities through unstructured data, such as free-text documents, and compare entities with no direct connections.
Innovation Solution
A system and technique that represent entities as vertices in a directed graph, using entity co-occurrences in unstructured documents and structured data sources to generate entity fingerprints, which are compared based on similarity scores combining supervised, unsupervised, and temporal factors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If entity-centric models rely on structured content, then data organization and retrieval are simplified, but the ability to represent entities through unstructured data and discover hidden connections is lost
Solution Approach 1:
The patent combines structured data sources (databases, spreadsheets) with unstructured data sources (text documents, web pages) into a unified entity-centric model. The system processes both structured and unstructured data together, extracting entity attributes from both sources and integrating them into a single comprehensive representation of each entity, thereby preserving information from both data types while maintaining organized access.
Solution Approach 2:
The patent introduces an intermediary processing layer that extracts entity mentions and attributes from unstructured text and bridges them to the structured entity model. This intermediary layer parses natural language text, identifies entity references, extracts relevant attributes, and links them to the corresponding structured entity records, enabling integration without requiring the unstructured data to conform to strict schemas.
2Ease of operation
If entities are compared using traditional attribute matching, then direct comparisons are straightforward, but entities with no direct connections cannot be meaningfully analyzed
Solution Approach 1:
The patent extends entity comparison from direct attribute matching to multi-hop relationship traversal in the entity graph. Instead of comparing only entities with direct attribute overlaps, the system navigates through the graph structure to discover indirect connections via intermediate entities, enabling comparison of entities that are not directly related by examining their shared neighbors and relationship patterns across multiple hops.
Solution Approach 2:
The patent segments the comparison process into distinct components: direct attribute comparison, relationship pattern analysis, and neighborhood similarity assessment. By breaking down the comparison task into these separate analytical layers, the system can evaluate both directly connected entities and indirectly related entities using appropriate methods for each type of relationship.
3Reliability
If the system processes both structured and unstructured data, then comprehensive entity representation is achieved, but system complexity increases
Solution Approach 1:
The patent implements self-service mechanisms where the system automatically discovers entity mentions in unstructured text, extracts relevant attributes, resolves entity identities against the structured database, and updates entity representations without manual intervention. The entity-centric model autonomously processes incoming data from multiple sources, performing entity resolution, attribute extraction, and relationship inference automatically, reducing the need for manual system configuration and data curation.
Data Source
AI summary
Systems and techniques for exploring relationships among entities are disclosed. The systems and techniques provide an entity-based information analysis and content aggregation platform that uses heterogeneous data sources to construct and maintain an ecosystem around tangible and logical entities. Entities are represented as vertices in a directed graph, and edges are generated using entity co-occurrences in unstructured documents and supervised information from structured data sources. Significance scores for the edges are computed using a method that combines supervised, unsupervised and temporal factors into a single score. Important entity attributes from the structured content and the entity neighborhood in the graph are automatically summarized as the entity fingerprint. Entities may be compared to one another based on similarity of their entity fingerprints. An interactive user interface is also disclosed that provides exploratory access to the graph and supports decision support processes.


