Entity Resolution via Learned Field Dependencies
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current database search methods for supplementing information about individuals are time-consuming and yield low-quality results due to the need for manual data entry and lack of understanding of data field dependencies between sources, leading to many irrelevant results.
Innovation Solution
A system and method for automatically generating queries across multiple data sources using machine learning techniques to match entity records, reducing user interaction and improving result relevance by leveraging dependencies between data fields.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If manual search methods are used to supplement information about individuals, then enterprises can access additional data from public databases, but the process becomes time-consuming and requires multiple manual tasks for each individual
Solution Approach 1:
The system enables automatic self-service by having the computer automatically generate search queries, execute searches across multiple data sources, and populate database fields without human intervention. The computer uses learned dependencies between data fields to autonomously supplement incomplete information.
Solution Approach 2:
The system performs preliminary action by pre-learning the dependencies between data fields from multiple data sources before actual information supplementation is needed. This pre-acquired knowledge is then applied to automatically generate appropriate search queries and map results to relevant fields.
2Loss of information
If manual search methods are used to supplement information about individuals, then enterprises can access additional data from public databases, but the quality of search results becomes very low due to lack of understanding of data field dependencies
Solution Approach 1:
The system implements feedback by learning from the relationships between data fields across multiple data sources. The computer analyzes how fields depend on each other and uses this learned feedback to generate more accurate search queries and better map results to the correct fields, improving result quality over time.
Solution Approach 2:
The system introduces an intermediary component that learns and understands the dependencies between data fields. This intermediary knowledge layer acts as a mediator between the search query generation and the data matching process, enabling more accurate interpretation and mapping of search results.
3Loss of information
If manual search methods are used to supplement information about individuals, then enterprises can access additional data from public databases, but the process requires an enterprise user to perform multiple manual tasks for each individual
Solution Approach 1:
The system enables automatic self-service by having the computer automatically generate search queries, execute searches across multiple data sources, and populate database fields without human intervention. The computer uses learned dependencies between data fields to autonomously supplement incomplete information.
Solution Approach 2:
The system replaces the mechanical manual operations of copying data, pasting into search fields, and manually evaluating results with an automated computer-based system. The mechanical tasks previously performed by users are substituted by algorithmic query generation and automatic result processing.
4Loss of information
If manual search methods are used to supplement information about individuals, then enterprises can access additional data from public databases, but without knowledge of data field dependencies, any search is suboptimal, yielding many irrelevant results
Solution Approach 1:
The system performs preliminary action by pre-learning the dependencies between data fields from multiple data sources before actual information supplementation is needed. This pre-acquired knowledge is then applied to automatically generate appropriate search queries and map results to relevant fields.
Solution Approach 2:
The system implements feedback by learning from the relationships between data fields across multiple data sources. The computer analyzes how fields depend on each other and uses this learned feedback to generate more accurate search queries and better map results to the correct fields, improving result quality over time.
Data Source
AI summary
Entity resolution techniques for matching entity records from different data sources are provided. In one technique, an entity record from a source database is identified along with multiple data items included therein. Each data item corresponds to an attribute of multiple source attributes. For one of the data items that corresponds to a first source attribute, multiple target attributes are identified. A first query is generated that includes the data items and associates the data item with each of the multiple target attributes. A second query that is different than the first query is also generated. Two searches are performed of a target database: one based on the first query and the other based on the second query. A scoring model generates multiple scores, one for each search result. It is determined whether the entity record matches an entity record in the target database based on the set of scores.


