Automated Record Linking Across Vendor Datasets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In Data as a Service (DaaS) systems, customers must manually identify similar records across separate datasets from different vendors, which can be time-consuming and inefficient, especially when dealing with datasets of different types and structures.
Innovation Solution
The DaaS system introduces a data assessment service that links records from separate vendor datasets based on similarity, using match keys and ingestion metadata to determine matching fields and generate linked records, thereby enhancing search and match query results without manual intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If customers manually identify similar records across separate vendor datasets, then they can find related information, but the process becomes time-consuming and inefficient
Solution Approach 1:
The system enables self-service by implementing an automated record linking mechanism that operates without customer intervention. The data assessment service automatically identifies and links similar records across vendor datasets using match keys and similarity algorithms, allowing the system to serve itself in performing the matching function that previously required manual customer effort.
Solution Approach 2:
The patent replaces the mechanical manual process of record matching with an automated computational system. The data assessment service uses algorithms that compare match keys, analyze data similarity, and automatically create links between records, substituting human manual operations with automated information processing mechanisms.
2Stability of the object's composition
If the DaaS system processes multiple vendor datasets separately, then each dataset maintains its structure, but customers cannot easily identify cross-dataset similar records
Solution Approach 1:
The patent introduces an intermediary mechanism in the form of match keys and a data assessment service that bridges separate vendor datasets. These match keys serve as intermediary identifiers that enable comparison and linking across datasets without altering the original dataset structures, facilitating cross-dataset record identification while preserving data integrity.
Solution Approach 2:
The data assessment service performs multiple functions: it maintains dataset structure integrity, identifies similar records across different vendors, links related records, and enhances query results. This multi-functional approach allows a single system component to address both structure preservation and ease of cross-dataset operation.
3Measurement precision
If the system provides separate query results for each dataset, then query accuracy is maintained, but customers must manually recognize similar records across results
Solution Approach 1:
The system performs preliminary action by automatically linking similar records across vendor datasets before the customer receives query results. The data assessment service pre-processes the data by creating relationships between similar records, so when customers receive query results, the similar records are already identified and linked, eliminating the need for manual recognition effort.
4Productivity
If the DaaS system implements automated record linking, then query results are enhanced, but system complexity increases
Solution Approach 1:
The patent segments the system into distinct functional components: vendor datasets, match keys, data assessment service, and query processing. This segmentation allows the automated record linking functionality to be added as a modular component without fundamentally complicating the entire system architecture, making the complexity manageable and localized to specific subsystems.
Data Source
AI summary
A method for linking records from different datasets based on record similarities is described. The method includes ingesting a first dataset, including a first set of records with a first set of fields, wherein the first dataset is associated with a first vendor and a first type of data, and a second dataset, including a second set of records with a second set of fields, wherein the second dataset is associated with a second vendor and a second type of data; determining that a first record from the first set of records is similar to a second record from the second set of records based on similarities between fields in the first and second set of fields; and linking the first and second records in response to determining that the similarity, wherein the first and second vendors are different and/or the first and second types of data are different.


