Entity Disambiguation via Latent Topic Graph and Centrality Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems for entity disambiguation face challenges such as entity recognition errors, ambiguous references, and inefficient processing due to exponential document growth, incomplete information, and redundancy, leading to computational bottlenecks and operational inefficiencies.

Innovation Solution

A system and method that generate an entity network by parsing documents to determine latent features, importance scores using centrality algorithms, relevance scores via link analysis, and clustering name entities, employing contextual information in a mathematical model for disambiguation, thereby reducing errors and computational complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional data mining and data scraping methods are used for entity disambiguation, then entity recognition can be performed, but entity recognition errors and ambiguous references occur due to incomplete, redundant or ambiguous information

Engineering Contradiction:
Improveentity recognition accuracyVSAvoidincomplete or ambiguous entity information
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent combines multiple information sources including text data, citation data, and metadata into a unified entity resolution framework. By merging these diverse data types, the system overcomes the limitations of individual data sources and reduces entity recognition errors caused by incomplete or ambiguous information in any single source.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces citation data as an intermediary element that connects entities across different documents. Citations serve as reliable mediators that provide verifiable relationships between entities, reducing ambiguous references and improving the accuracy of entity disambiguation by providing external validation beyond the primary text.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If documents are stored as a large collection of objects with exponential increase in number, then comprehensive information is available, but computational performance becomes a bottleneck for entity extraction and retrieval

Engineering Contradiction:
Improvenumber of documentsVSAvoidcomputational performance
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent segments the large collection of documents into manageable units by extracting and processing individual entities, their attributes, and relationships separately. This segmentation allows the system to handle exponential document growth efficiently by focusing computational resources on specific entity extractions rather than processing entire document collections at once.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary entity extraction and normalization during the document ingestion phase, creating structured entity data before retrieval operations. By pre-processing and organizing entity information in advance, the system reduces computational overhead during entity extraction and retrieval operations, maintaining productivity despite exponential document growth.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If multiple duplicate entities representing the same entity are present in documents, then comprehensive coverage is achieved, but entity extraction and retrieval efficiency is greatly affected

Engineering Contradiction:
Improveentity coverageVSAvoidentity extraction time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent implements feedback mechanisms that use entity similarity metrics and relationship patterns to identify and merge duplicate entities. By continuously evaluating entity relationships and providing feedback on entity equivalence, the system resolves duplicates efficiently while maintaining comprehensive coverage, reducing entity extraction time without losing entity diversity.

Inventive Principle:
Principle #23Feedback

4Ease of operation

If conventional approaches are used for text normalization with multiple entity types, then processing can be performed, but the process becomes difficult to execute and cumbersome

Engineering Contradiction:
Improvetext normalization processVSAvoidnormalization system complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent implements a universal entity normalization framework that handles multiple entity types (persons, organizations, locations, etc.) through a single unified process. This multi-functional approach simplifies operation by providing consistent normalization procedures across different entity types, reducing the complexity and difficulty of executing text normalization despite the diversity of entity types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11243985B1System and method for name entity disambiguation with latent topic and deep graph analysis
Publication Date: 2022.02.08 INNOPLEXUS AG
  • US11243985B1 patent drawing
  • US11243985B1 patent drawing
  • US11243985B1 patent drawing

AI summary

A system for entity disambiguation. The system includes a server arrangement communicably coupled to a database arrangement including a plurality of documents. The server arrangement is configured to generate an entity network by parsing the plurality of documents, determining latent features relating to the plurality of documents and identifying relationships between the latent features and at least one document entity; determine importance score of each of the name entities using centrality algorithms; determine a relevance score of at least one document entity using link analysis algorithms; cluster the name entities based on the relationships therebetween using clustering coefficients; assign a name entity from each cluster of name entities as the person entity; and employ contextual information determined for the entity network to be used in a mathematical model for entity disambiguation.