Iterative Entity Resolution for Accurate NLP Knowledge Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing entity analysis processes are time-consuming, prone to errors, and lack standardization due to human variability, making them inefficient and difficult to audit.
Innovation Solution
A computer-implemented system using machine-learning models iteratively classifies candidate entities in content pages, analyzing similarities with confirmed data to highlight confirmed entities, reducing human error and improving processing speed and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If human researchers manually search and analyze content pages to identify target entities, then entity resolution can be performed with human judgment and adaptability, but the process becomes time-consuming, resource-intensive, and prone to errors
Solution Approach 1:
The patent replaces the mechanical system of manual human research with an automated computing device system that performs entity resolution through machine learning models and natural language processing, eliminating the need for human researchers to manually search and analyze content pages while maintaining or improving accuracy
Solution Approach 2:
The system enables self-service entity resolution by automatically collecting content pages, extracting candidate entities, comparing them against target entities, and generating resolution results without human intervention, allowing the system to serve itself in completing the entire entity resolution workflow
2Adaptability or versatility
If human researchers perform entity analysis processes, then flexible judgment and adaptability can be applied, but the processes cannot be audited and are not reproducible
Solution Approach 1:
The patent replaces human researcher judgment with automated machine learning models that provide consistent, reproducible, and auditable entity resolution results through deterministic algorithms while maintaining adaptability through trained models that can handle various entity types and contexts
3Ease of manufacture
If a typical computing device is used to extract knowledge from unstructured content pages, then hardware costs can be kept low, but the process becomes impractical and inefficient
Solution Approach 1:
The patent replaces inefficient manual or basic automated text processing with advanced machine learning-based natural language processing capabilities that can efficiently extract and compare entity information from unstructured content pages, dramatically improving productivity while remaining implementable on standard computing devices
Data Source
AI summary
A computer-implemented system to perform group entity identification, resolution and knowledge extraction is provided. The system receives an indication of one or more potentially related entities and basic attributes. The system then collects a plurality of content pages comprising candidate attribute data related to one or more candidate entities. Based on entity resolution configuration and entity-resolution module which employs deep-learning models, the system obtains initial additional confirmed entity attribute data or relevant attribute data. With additional knowledges acquired, the system iteratively goes over the same contents again and potentially classifies entities identified in the content pages to be at least confirmed, relevant, or irrelevant entities, until no more additional confident knowledges obtained for target entities during iteration. After iterations of entity resolution processes, the system finally extracts entity knowledge based on predefined knowledge map for individuals and business entities, summarization of knowledges for entities are then performed, results are displayed.


