Feature Vector Generation for Probabilistic Record Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current Master Data Management (MDM) solutions for record matching and linking across different sources are inefficient, requiring manual configuration and expertise, leading to potential errors and time-consuming processes.

Innovation Solution

A computer-implemented method that automatically generates feature vectors for record attributes, uses machine learning to determine similarity scores, and links records based on a confidence threshold, thereby streamlining the record matching process across different sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual configuration and expertise are used for record matching, then matching accuracy can be maintained, but the process becomes time-consuming and error-prone

Engineering Contradiction:
Improvematching accuracyVSAvoidtime consumption
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs self-service by automatically generating feature vectors and configuring matching parameters without requiring manual expert intervention. The machine learning model autonomously processes record attributes, generates similarity scores, and produces matching results, eliminating the need for manual configuration while maintaining matching accuracy.

Inventive Principle:
Principle #25Self-service

2Manufacturing precision

If manual expert configuration is used for feature vector generation, then matching quality is preserved, but human error and time consumption increase

Engineering Contradiction:
Improvematching qualityVSAvoidhuman error
Core Design Contradiction:
Manufacturing precisionVSObject-generated harmful factors

Solution Approach 1:

The patent replaces the mechanical system of manual expert configuration with an automated machine learning system. The machine learning model processes record attributes, generates feature vectors, and computes similarity scores automatically, substituting human manual operations with an automated computational system that eliminates human error while preserving matching quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If automated feature vector generation is implemented, then efficiency and productivity improve, but system complexity increases

Engineering Contradiction:
Improvematching efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces feature vectors as an intermediary representation between raw record attributes and matching decisions. The machine learning model transforms complex record data into standardized feature vectors that capture essential characteristics, simplifying the subsequent matching process while improving efficiency. This intermediary layer manages system complexity by providing a structured intermediate representation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12039273B2Feature vector generation for probabalistic matching
Publication Date: 2024.07.16 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12039273B2 patent drawing
  • US12039273B2 patent drawing
  • US12039273B2 patent drawing

AI summary

A computer-implemented method increases the efficiency of matching records from two sources. The method includes identifying a first source and a second source wherein each of the sources include one or more records and each record includes one or more attributes. The method further includes determining, based on a corpus, the one or more attributes and generating, based on the attributes, a set of feature vectors which vectors represent the one or more attributes. The method includes comparing each record in the first source against each record in the second source. The method further includes generating, in response to the comparing, a link confidence. The method also includes linking, in response to the link confidence being above a linking threshold, the associated records. The method includes determining a first feature vector of the set of feature vectors used in the linking, and outputting a set of results.