Feature Vector Generation for Probabilistic Record Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current Master Data Management (MDM) solutions for record matching and linking across different sources are inefficient, requiring manual configuration and expertise, leading to potential errors and time-consuming processes.
Innovation Solution
A computer-implemented method that automatically generates feature vectors for record attributes, uses machine learning to determine similarity scores, and links records based on a confidence threshold, thereby streamlining the record matching process across different sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual configuration and expertise are used for record matching, then matching accuracy can be maintained, but the process becomes time-consuming and error-prone
Solution Approach 1:
The system performs self-service by automatically generating feature vectors and configuring matching parameters without requiring manual expert intervention. The machine learning model autonomously processes record attributes, generates similarity scores, and produces matching results, eliminating the need for manual configuration while maintaining matching accuracy.
2Manufacturing precision
If manual expert configuration is used for feature vector generation, then matching quality is preserved, but human error and time consumption increase
Solution Approach 1:
The patent replaces the mechanical system of manual expert configuration with an automated machine learning system. The machine learning model processes record attributes, generates feature vectors, and computes similarity scores automatically, substituting human manual operations with an automated computational system that eliminates human error while preserving matching quality.
3Productivity
If automated feature vector generation is implemented, then efficiency and productivity improve, but system complexity increases
Solution Approach 1:
The patent introduces feature vectors as an intermediary representation between raw record attributes and matching decisions. The machine learning model transforms complex record data into standardized feature vectors that capture essential characteristics, simplifying the subsequent matching process while improving efficiency. This intermediary layer manages system complexity by providing a structured intermediate representation.
Data Source
AI summary
A computer-implemented method increases the efficiency of matching records from two sources. The method includes identifying a first source and a second source wherein each of the sources include one or more records and each record includes one or more attributes. The method further includes determining, based on a corpus, the one or more attributes and generating, based on the attributes, a set of feature vectors which vectors represent the one or more attributes. The method includes comparing each record in the first source against each record in the second source. The method further includes generating, in response to the comparing, a link confidence. The method also includes linking, in response to the link confidence being above a linking threshold, the associated records. The method includes determining a first feature vector of the set of feature vectors used in the linking, and outputting a set of results.


