Entity Disambiguation via Latent Topic Graph and Centrality Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems for entity disambiguation face challenges such as entity recognition errors, ambiguous references, and inefficient processing due to exponential document growth, incomplete information, and redundancy, leading to computational bottlenecks and operational inefficiencies.
Innovation Solution
A system and method that generate an entity network by parsing documents to determine latent features, importance scores using centrality algorithms, relevance scores via link analysis, and clustering name entities, employing contextual information in a mathematical model for disambiguation, thereby reducing errors and computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional data mining and data scraping methods are used for entity disambiguation, then entity recognition can be performed, but entity recognition errors and ambiguous references occur due to incomplete, redundant or ambiguous information
Solution Approach 1:
The patent combines multiple information sources including text data, citation data, and metadata into a unified entity resolution framework. By merging these diverse data types, the system overcomes the limitations of individual data sources and reduces entity recognition errors caused by incomplete or ambiguous information in any single source.
Solution Approach 2:
The patent introduces citation data as an intermediary element that connects entities across different documents. Citations serve as reliable mediators that provide verifiable relationships between entities, reducing ambiguous references and improving the accuracy of entity disambiguation by providing external validation beyond the primary text.
2Quantity of substance
If documents are stored as a large collection of objects with exponential increase in number, then comprehensive information is available, but computational performance becomes a bottleneck for entity extraction and retrieval
Solution Approach 1:
The patent segments the large collection of documents into manageable units by extracting and processing individual entities, their attributes, and relationships separately. This segmentation allows the system to handle exponential document growth efficiently by focusing computational resources on specific entity extractions rather than processing entire document collections at once.
Solution Approach 2:
The patent performs preliminary entity extraction and normalization during the document ingestion phase, creating structured entity data before retrieval operations. By pre-processing and organizing entity information in advance, the system reduces computational overhead during entity extraction and retrieval operations, maintaining productivity despite exponential document growth.
3Adaptability or versatility
If multiple duplicate entities representing the same entity are present in documents, then comprehensive coverage is achieved, but entity extraction and retrieval efficiency is greatly affected
Solution Approach 1:
The patent implements feedback mechanisms that use entity similarity metrics and relationship patterns to identify and merge duplicate entities. By continuously evaluating entity relationships and providing feedback on entity equivalence, the system resolves duplicates efficiently while maintaining comprehensive coverage, reducing entity extraction time without losing entity diversity.
4Ease of operation
If conventional approaches are used for text normalization with multiple entity types, then processing can be performed, but the process becomes difficult to execute and cumbersome
Solution Approach 1:
The patent implements a universal entity normalization framework that handles multiple entity types (persons, organizations, locations, etc.) through a single unified process. This multi-functional approach simplifies operation by providing consistent normalization procedures across different entity types, reducing the complexity and difficulty of executing text normalization despite the diversity of entity types.
Data Source
AI summary
A system for entity disambiguation. The system includes a server arrangement communicably coupled to a database arrangement including a plurality of documents. The server arrangement is configured to generate an entity network by parsing the plurality of documents, determining latent features relating to the plurality of documents and identifying relationships between the latent features and at least one document entity; determine importance score of each of the name entities using centrality algorithms; determine a relevance score of at least one document entity using link analysis algorithms; cluster the name entities based on the relationships therebetween using clustering coefficients; assign a name entity from each cluster of name entities as the person entity; and employ contextual information determined for the entity network to be used in a mathematical model for entity disambiguation.


