Entity Resolution via Learned Field Dependencies

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current database search methods for supplementing information about individuals are time-consuming and yield low-quality results due to the need for manual data entry and lack of understanding of data field dependencies between sources, leading to many irrelevant results.

Innovation Solution

A system and method for automatically generating queries across multiple data sources using machine learning techniques to match entity records, reducing user interaction and improving result relevance by leveraging dependencies between data fields.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If manual search methods are used to supplement information about individuals, then enterprises can access additional data from public databases, but the process becomes time-consuming and requires multiple manual tasks for each individual

Engineering Contradiction:
Improvecompleteness of individual informationVSAvoidtime required for manual search
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system enables automatic self-service by having the computer automatically generate search queries, execute searches across multiple data sources, and populate database fields without human intervention. The computer uses learned dependencies between data fields to autonomously supplement incomplete information.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary action by pre-learning the dependencies between data fields from multiple data sources before actual information supplementation is needed. This pre-acquired knowledge is then applied to automatically generate appropriate search queries and map results to relevant fields.

Inventive Principle:
Principle #10Preliminary action

2Loss of information

If manual search methods are used to supplement information about individuals, then enterprises can access additional data from public databases, but the quality of search results becomes very low due to lack of understanding of data field dependencies

Engineering Contradiction:
Improvecompleteness of individual informationVSAvoidaccuracy of search results
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The system implements feedback by learning from the relationships between data fields across multiple data sources. The computer analyzes how fields depend on each other and uses this learned feedback to generate more accurate search queries and better map results to the correct fields, improving result quality over time.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system introduces an intermediary component that learns and understands the dependencies between data fields. This intermediary knowledge layer acts as a mediator between the search query generation and the data matching process, enabling more accurate interpretation and mapping of search results.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If manual search methods are used to supplement information about individuals, then enterprises can access additional data from public databases, but the process requires an enterprise user to perform multiple manual tasks for each individual

Engineering Contradiction:
Improvecompleteness of individual informationVSAvoidease of information supplementation
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system enables automatic self-service by having the computer automatically generate search queries, execute searches across multiple data sources, and populate database fields without human intervention. The computer uses learned dependencies between data fields to autonomously supplement incomplete information.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system replaces the mechanical manual operations of copying data, pasting into search fields, and manually evaluating results with an automated computer-based system. The mechanical tasks previously performed by users are substituted by algorithmic query generation and automatic result processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Loss of information

If manual search methods are used to supplement information about individuals, then enterprises can access additional data from public databases, but without knowledge of data field dependencies, any search is suboptimal, yielding many irrelevant results

Engineering Contradiction:
Improvecompleteness of individual informationVSAvoideffectiveness of information supplementation
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system performs preliminary action by pre-learning the dependencies between data fields from multiple data sources before actual information supplementation is needed. This pre-acquired knowledge is then applied to automatically generate appropriate search queries and map results to relevant fields.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback by learning from the relationships between data fields across multiple data sources. The computer analyzes how fields depend on each other and uses this learned feedback to generate more accurate search queries and better map results to the correct fields, improving result quality over time.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11080272B2Entity resolution techniques for matching entity records from different data sources
Publication Date: 2021.08.03 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11080272B2 patent drawing
  • US11080272B2 patent drawing
  • US11080272B2 patent drawing

AI summary

Entity resolution techniques for matching entity records from different data sources are provided. In one technique, an entity record from a source database is identified along with multiple data items included therein. Each data item corresponds to an attribute of multiple source attributes. For one of the data items that corresponds to a first source attribute, multiple target attributes are identified. A first query is generated that includes the data items and associates the data item with each of the multiple target attributes. A second query that is different than the first query is also generated. Two searches are performed of a target database: one based on the first query and the other based on the second query. A scoring model generates multiple scores, one for each search result. It is determined whether the entity record matches an entity record in the target database based on the set of scores.