Dynamic Data Record Linking with Categorical Similarity Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data mining systems struggle to effectively link incongruous data records describing the same individual due to inconsistencies and errors, making it challenging to correlate and combine diverse sets of data records accurately.

Innovation Solution

A method and system that parses and categorizes data records using dynamic rules, calculates similarity scores between categories, and modifies identities based on these scores to link records when they exceed a threshold, utilizing a computing device with modules for parsing, matching, and control to manage processing needs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data records are collected from multiple sources over time, then the quantity of available data increases, but the consistency and accuracy of the data decreases due to updates and errors

Engineering Contradiction:
Improvequantity of data recordsVSAvoidconsistency and accuracy of data
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system performs preliminary parsing and categorization of data records before linking them. By pre-processing the data to identify and standardize categories (e.g., name, address, date of birth), the system prepares the data in advance for accurate matching, reducing the impact of inconsistencies that arise from multiple updates and sources

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary similarity score calculation mechanism that acts as a mediator between raw data records and the linking decision. This intermediary layer compares categorized information from different records and determines similarity scores, which then guide the linking process, filtering out erroneous or inconsistent data while preserving accurate connections

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If contact information is updated frequently to reflect current data, then the data remains current and accurate, but the risk of inconsistencies and entry errors increases

Engineering Contradiction:
Improvecurrent accuracy of dataVSAvoiddata entry accuracy
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The system uses feedback from similarity score calculations to verify and correct data entries. When records are linked based on initial parsing, the system can cross-validate the linked information and provide feedback to correct potential entry errors, ensuring that updated contact information maintains both current accuracy and consistency across records

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent transforms raw data parameters into standardized categories through parsing. By changing the representation of data from raw formats to standardized categories (e.g., converting various address formats to a standard structure), the system maintains data currentness while reducing the impact of entry errors through standardization

Inventive Principle:
Principle #35Parameter changes

3Productivity

If data records are parsed and categorized using static rules, then the parsing process is simple and fast, but the system cannot adapt to varying data formats and structures

Engineering Contradiction:
Improveparsing speedVSAvoidadaptability to data formats
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic parsing rules that can adapt to varying data formats. Rather than using fixed static rules, the system dynamically determines how to parse and categorize information based on the structure and content of each data record, enabling both speed and adaptability in processing diverse data formats from multiple sources

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250231966A1Managing data processing efficiency, and applications thereof
Publication Date: 2025.07.17 H1 INSIGHTS INC
  • US20250231966A1 patent drawing
  • US20250231966A1 patent drawing
  • US20250231966A1 patent drawing

AI summary

Disclosed herein are system, method, and computer program product embodiments for linking data records in memory. The system, method, and computer program product includes accessing a first record stored in memory, the first record holding information describing a first person and accessing at least one additional record stored in memory, the additional records holding information describing additional persons. The method continues by parsing the information of the first record and additional record and assigning the parsed information to predefined categories within the respective records. After assigning the information into categories, a similarity score between categorical information in the first record and categorical information of additional records is determined. A category of an additional record is then modified based on the similarity score, so the additional record is associated with the first person.