Entity Resolution Metric for Mismatched Attribute Cardinalities
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional algorithms for determining association metrics between records from different data sources with mismatched attributes, such as those with the same data type but different meanings or disparate data collection methods, often fail to accurately correlate records associated with the same entity.
Innovation Solution
A system that compares individual association metrics between mismatched attributes, applying two levels of reduction to generate a mismatched-attribute association metric, which is then used to train and apply an entity resolution model using machine learning, enabling the determination of record associations across different schemas and data sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional algorithms are used to determine association metrics between attributes from different schemas, then the processing is simple and fast, but the accuracy of identifying records associated with the same entity deteriorates when attributes have different meanings or cardinalities
Solution Approach 1:
The association metric determination is segmented into two levels: individual association metrics for each attribute pair, and then a reduced association metric that aggregates these individual metrics. This segmentation allows the system to handle complex mismatched attributes by breaking down the problem into manageable components while maintaining high accuracy.
Solution Approach 2:
The patent creates a composite association metric by combining multiple individual association metrics through a reduction process. This composite metric integrates information from multiple attribute comparisons, enabling accurate entity resolution even when individual attributes have different meanings or cardinalities across schemas.
2Measurement precision
If attributes from different schemas with different cardinalities are compared directly, then the processing is straightforward, but the association metric accuracy deteriorates
Solution Approach 1:
The patent introduces a new dimension of analysis by creating individual association metrics for each attribute pair before aggregating them into a reduced association metric. This two-level approach adds a dimensional layer that captures the nuanced relationships between mismatched attributes, enabling accurate comparison even when cardinalities differ.
Solution Approach 2:
The system changes the parameters of comparison by transforming raw attribute values into standardized association metrics that account for different cardinalities and meanings. The reduction process further transforms these individual metrics into a comprehensive reduced association metric, adapting the comparison parameters to handle schema mismatches effectively.
3Reliability
If data sources with different collection methods are integrated, then the data coverage is comprehensive, but the consistency and reliability of association metrics deteriorate
Solution Approach 1:
The individual association metrics and the reduced association metric serve as intermediaries between data from different sources with different collection methods. These metrics mediate the comparison process, translating disparate data formats and collection methodologies into a unified framework that maintains reliability across diverse data sources.
Solution Approach 2:
The reduced association metric provides a universal measure that works across multiple data sources and collection methods. By aggregating individual association metrics, the system creates a multi-functional metric that can reliably assess record associations regardless of the original data source or collection methodology, enabling comprehensive data integration.
Data Source
AI summary
An association metric for record attributes associated with cardinalities that are not necessarily the same is used for training and/or applying an entity resolution (ER) model. A pair of records includes (a) a first record indicating a first set of values for a first attribute and (b) a second record indicating a second set of values for a second attribute. Each of the first set of values and each of the second set of values are compared to determine individual association metrics. A first-level reduction operation is applied to subsets of the individual association metrics to determine reduced association metrics. A second-level reduction operation is applied to the reduced association metrics to determine an association metric, for the pair of records, for training and/or applying an ER model.


