Machine Learning Record Matching with Pre-Cleaning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing record matching techniques struggle with accuracy and efficiency due to inconsistent data formats and inaccurate information across different databases, leading to errors and inefficiencies in information sharing.

Innovation Solution

The method employs multiple machine learning models for pre-cleaning records, computing relevance scores, and identifying matching records, utilizing techniques such as BERT NER for data validation and ranking algorithms for weight computation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional record matching techniques are used, then the matching process can be performed, but the accuracy is low and errors occur due to inconsistent data formats and inaccurate information

Engineering Contradiction:
Improverecord matching accuracyVSAvoidmatching reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent applies preliminary action by pre-cleaning records before matching using machine learning models. The system identifies and corrects inconsistent data formats, fills missing values, and validates information quality before the actual matching process, thereby improving both accuracy and reliability of record matching results

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces machine learning models as intermediaries between the raw records and the matching process. These models act as mediators that transform inconsistent, unstructured data into clean, standardized formats that can be reliably compared, thus resolving the contradiction between handling diverse data formats and maintaining matching accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If traditional record matching techniques are used, then the process can complete, but it is slow and computationally expensive

Engineering Contradiction:
Improverecord matching speedVSAvoidcomputational resource consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent performs record cleaning and validation in advance using efficient machine learning models, reducing the complexity of the subsequent matching process. By preprocessing records to standardize formats and fill missing values beforehand, the actual matching operation becomes faster and requires fewer computational resources

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces traditional mechanical record matching approaches with machine learning-based systems. The ML models automatically learn patterns and relationships in the data, substituting manual or rule-based matching mechanisms with intelligent algorithms that are both faster and more resource-efficient

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20250077528A1Fast record matching using machine learning
Publication Date: 2025.03.06 INTUIT INC
  • US20250077528A1 patent drawing
  • US20250077528A1 patent drawing
  • US20250077528A1 patent drawing

AI summary

The present disclosure provides techniques for fast record matching using machine learning. One example method includes receiving a request indicating one or more attributes, identifying, from a plurality of records using a first machine learning model, a set of records, wherein each record of the set of records indicates the one or more attributes, computing, for each record of the set of records using a second machine learning model, a first relevance score for the record, computing, for each record of the set of records using a third machine learning model, a second relevance score for the record, and identifying, based on the first relevance score for each record of the set of records and the second relevance score for each record of the set of records, a given record of the set of records best matching the request.