Cascade Record Matching System for Unstructured Data Integration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data integration methods fail to effectively organize and integrate data without common identifiers, especially when data is unstructured, incomplete, or contains errors, leading to inefficiencies in record linkage and data management.

Innovation Solution

A system and method for matching and assembling data records that can handle unstructured or incomplete data by creating new records through a process of matching, grouping, and inferring additional information, even in the absence of common identifiers, using a cascade of processing steps to refine matches and assemble multidimensional and heterogeneous records.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional record linkage techniques are used that depend on common explicit identifiers, then matching accuracy for records with identifiers is improved, but the system cannot effectively match records without common identifiers or with implicit entities

Engineering Contradiction:
Improvematching accuracyVSAvoidcapability to match records without common identifiers
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary identifier generation mechanism that creates synthetic identifiers from unstructured data fields. Instead of requiring pre-existing common identifiers, the system extracts features from data fields, generates candidate identifiers, and uses them as intermediaries to enable matching between records that would otherwise be incompatible. This resolves the contradiction by maintaining matching accuracy through structured identifier comparison while extending adaptability to handle unstructured and identifier-less records.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system dynamically changes the parameters used for matching by generating multiple candidate identifiers from different data field combinations and evaluating them against threshold criteria. Rather than relying on a fixed identifier schema, the system adapts its matching parameters based on the structure and quality of the input data, allowing it to effectively match records with implicit entities while maintaining precision through configurable threshold-based filtering.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If data is unstructured or contains errors, then data availability and coverage are improved, but data quality and reliability deteriorate

Engineering Contradiction:
Improvedata coverageVSAvoiddata quality
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system performs preliminary data validation and error detection before the main matching process. It extracts data fields, validates their structure, and identifies potential errors or inconsistencies upfront. This preliminary action allows the system to handle unstructured and error-prone data by preprocessing it into a standardized format, thereby maintaining high data coverage while ensuring reliability through early error detection and correction mechanisms.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback loops where matching results are evaluated against quality thresholds, and unsuccessful matches or low-confidence results trigger re-processing with adjusted parameters. The system uses feedback from the matching process to refine identifier generation, adjust validation criteria, and improve overall data quality. This feedback mechanism enables the system to maintain high data coverage while continuously improving reliability through iterative refinement.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If a cascade of processing steps is used to refine matches, then matching precision is improved, but processing complexity and time increase

Engineering Contradiction:
Improvematching precisionVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the matching process into distinct modular stages: data field extraction, candidate identifier generation, threshold-based filtering, and final matching. Each stage operates independently with well-defined inputs and outputs, allowing the system to achieve high matching precision through multiple refinement steps while managing complexity through modular design. This segmentation enables the cascade of processing steps to be implemented systematically without overwhelming system complexity.

Inventive Principle:
Principle #1Segmentation

4Adaptability or versatility

If multiple data sources with different structures are integrated, then data versatility and coverage are improved, but integration difficulty and heterogeneity management increase

Engineering Contradiction:
Improvedata source compatibilityVSAvoidintegration complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements a universal data field extraction framework that can handle multiple data source structures through a common interface. It uses configurable field identification rules and flexible data mapping capabilities to accommodate heterogeneous data formats while maintaining a unified processing pipeline. This universality allows the system to integrate diverse data sources with different structures, improving versatility while managing integration complexity through standardized processing mechanisms.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8166033B2System and method for matching and assembling records
Publication Date: 2012.04.24 ELSEVIER INC
  • US8166033B2 patent drawing
  • US8166033B2 patent drawing
  • US8166033B2 patent drawing

AI summary

A system and method for matching and assembling records is provided. One embodiment of the invention assembles records by applying a method for grouping records based on matching fields, assembling a new record as a composite of the matched records, and then repeating the grouping, matching and assembling steps in a cascade where the matching grouping and assembling steps are modified as a function of the cascade step and the assembled records created in earlier steps.