Record Linkage via Distance Measures and Standardization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data record linkage systems struggle to accurately associate records from multiple sources due to variations in data entry, formatting, and syntax, leading to incomplete or incorrect data retrieval, especially in online databases and apps, which limits the ability to find comprehensive and reliable information about entities.
Innovation Solution
The development of improved deduplication and linkage techniques that utilize distance measures and standardization methods to compare records, calculate quality measures for database reliability, and merge linked records into a unified database, ensuring more complete and accurate data presentation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional record linkage systems are used to associate records from multiple sources, then the system structure is simple, but the accuracy of record association deteriorates due to variations in data entry, formatting, and syntax
Solution Approach 1:
The patent applies preliminary action by standardizing data attributes before comparison. The system transforms raw data from multiple sources into a standardized format using predefined schemas and data transformation rules, which prepares the data for accurate matching. This preliminary standardization step enables conventional linkage systems to achieve higher accuracy without requiring fundamentally complex new architectures.
Solution Approach 2:
The patent introduces an intermediary layer between data sources and the linkage system. This intermediary includes data transformation modules, schema mapping components, and normalization functions that convert diverse data formats into a unified structure. This intermediary layer absorbs the complexity of handling variations in data entry, formatting, and syntax, allowing the core linkage system to operate with improved accuracy.
2Loss of information
If data records from multiple sources are linked without standardization, then the processing time is short, but the completeness of data retrieval deteriorates due to data fragmentation and inaccessibility
Solution Approach 1:
The system performs preliminary data standardization and schema alignment before the actual linkage operation. By pre-processing data to establish consistent formats and structures, the system reduces information loss during association while minimizing the time penalty through efficient preprocessing algorithms and cached transformation rules.
Solution Approach 2:
The patent segments the data linkage process into distinct phases: data ingestion, standardization, matching, and consolidation. This segmentation allows the system to handle large volumes of fragmented data from multiple sources systematically, reducing information loss by ensuring each segment is properly standardized before integration, while managing processing time through parallel processing of independent data segments.
3Reliability
If multiple data records for the same entity are created from separate information sources, then the quantity of data is increased, but the reliability of data association deteriorates due to duplicate records with different names or identification numbers
Solution Approach 1:
The patent introduces intermediary components that act as a buffer between raw data ingestion and final record association. These intermediaries include data validation layers, duplicate detection algorithms, and confidence scoring mechanisms that evaluate the reliability of associations. By processing records through this intermediary layer, the system can handle large quantities of data from multiple sources while maintaining high association reliability through systematic filtering and verification.
4Measurement precision
If conventional query methods are used to retrieve data from databases, then the ease of operation is high, but the accuracy of entity identification deteriorates when multiple identical-looking records exist
Solution Approach 1:
The patent implements feedback mechanisms in the query system that provide confidence scores and match quality metrics to users. When multiple records are found, the system returns them ranked by association confidence, with feedback indicators showing the reliability of each match. This allows users to easily identify the most accurate results without complex manual verification, maintaining ease of operation while dramatically improving entity identification accuracy.
Data Source
AI summary
Embodiments of methods for record linkage and comparing attributes are presented herein. Broadly speaking, embodiments of the present invention associate data records using distance measures and weights. More particularly, embodiments of the present invention generate a weight-based comparison of attributes. More specifically, embodiments of the present invention involve deduplication of records in a single database and linkage of records in multiple databases. In some embodiments of the invention, records may be merged into a master database based on the linkage and comparison of attributes. In addition, embodiments of the present invention may calculate a quality measure to be used in comparing attributes and records.


