Dynamic Data Record Linking with Categorical Similarity Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data mining systems struggle to effectively link incongruous data records describing the same individual due to inconsistencies and errors, making it challenging to correlate and combine diverse sets of data records accurately.
Innovation Solution
A method and system that parses and categorizes data records using dynamic rules, calculates similarity scores between categories, and modifies identities based on these scores to link records when they exceed a threshold, utilizing a computing device with modules for parsing, matching, and control to manage processing needs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data records are collected from multiple sources over time, then the quantity of available data increases, but the consistency and accuracy of the data decreases due to updates and errors
Solution Approach 1:
The system performs preliminary parsing and categorization of data records before linking them. By pre-processing the data to identify and standardize categories (e.g., name, address, date of birth), the system prepares the data in advance for accurate matching, reducing the impact of inconsistencies that arise from multiple updates and sources
Solution Approach 2:
The patent introduces an intermediary similarity score calculation mechanism that acts as a mediator between raw data records and the linking decision. This intermediary layer compares categorized information from different records and determines similarity scores, which then guide the linking process, filtering out erroneous or inconsistent data while preserving accurate connections
2Reliability
If contact information is updated frequently to reflect current data, then the data remains current and accurate, but the risk of inconsistencies and entry errors increases
Solution Approach 1:
The system uses feedback from similarity score calculations to verify and correct data entries. When records are linked based on initial parsing, the system can cross-validate the linked information and provide feedback to correct potential entry errors, ensuring that updated contact information maintains both current accuracy and consistency across records
Solution Approach 2:
The patent transforms raw data parameters into standardized categories through parsing. By changing the representation of data from raw formats to standardized categories (e.g., converting various address formats to a standard structure), the system maintains data currentness while reducing the impact of entry errors through standardization
3Productivity
If data records are parsed and categorized using static rules, then the parsing process is simple and fast, but the system cannot adapt to varying data formats and structures
Solution Approach 1:
The patent implements dynamic parsing rules that can adapt to varying data formats. Rather than using fixed static rules, the system dynamically determines how to parse and categorize information based on the structure and content of each data record, enabling both speed and adaptability in processing diverse data formats from multiple sources
Data Source
AI summary
Disclosed herein are system, method, and computer program product embodiments for linking data records in memory. The system, method, and computer program product includes accessing a first record stored in memory, the first record holding information describing a first person and accessing at least one additional record stored in memory, the additional records holding information describing additional persons. The method continues by parsing the information of the first record and additional record and assigning the parsed information to predefined categories within the respective records. After assigning the information into categories, a similarity score between categorical information in the first record and categorical information of additional records is determined. A category of an additional record is then modified based on the similarity score, so the additional record is associated with the first person.


