Active Entity Resolution Model for Procurement Data Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Duplicate data records in procurement and supply chain systems cause data messiness and incompleteness, making it difficult to search for suppliers and source goods or services efficiently.
Innovation Solution
An Active Entity Resolution (AER) model system that detects duplicate records, clusters them, generates canonical records, and uses machine learning to provide recommendations for new data records by matching them with master data, thereby improving data integrity and search functionality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If duplicate data records are allowed to exist in the system, then the quantity of data records increases, but data quality and search efficiency deteriorate
Solution Approach 1:
The system extracts duplicate records from the data set using entity resolution algorithms. The AER model identifies and separates duplicate supplier records from unique records, allowing the system to maintain a complete data set while eliminating the harmful effects of duplication through selective extraction of duplicate instances.
Solution Approach 2:
The system changes the state of data records by assigning confidence scores and similarity metrics to identify duplicates. By transforming the data through machine learning models that calculate match probabilities, the system can distinguish between unique and duplicate records without deleting any data, thus maintaining quantity while improving quality.
2Reliability
If manual processes are used to identify and manage duplicate records, then data accuracy can be maintained, but time consumption and operational complexity increase
Solution Approach 1:
The system performs self-service by automatically detecting, clustering, and managing duplicate records without human intervention. The AER model autonomously processes data records, identifies duplicates through machine learning, and generates recommendations, eliminating the need for manual data cleaning while maintaining high accuracy through algorithmic consistency.
Solution Approach 2:
The patent replaces manual mechanical processes with an automated machine learning system. The AER model uses machine learning algorithms to substitute human analysts, automatically comparing records, calculating similarity scores, and identifying duplicates, thereby maintaining data accuracy while dramatically reducing time consumption.
3Device complexity
If traditional entity resolution methods are used, then implementation simplicity is maintained, but resolution accuracy and automation capability deteriorate
Solution Approach 1:
The AER model serves multiple functions within a single unified system: it performs entity resolution, duplicate detection, data clustering, and recommendation generation. This multi-functional approach maintains relative system simplicity while achieving high automation capability, as the single model handles various data quality tasks that would otherwise require multiple separate tools.
Data Source
AI summary
Systems and methods are provided for accessing master data comprising a plurality of representative data records, where each representative data record represents a cluster of similar data records, and each similar data record has a confidence score indicating a confidence level that the similar data record corresponds to the cluster, and comparing a new data record to each representative data record of the plurality of representative data records using a machine learning model to generate a distance score. The systems and methods further provide for analyzing the cluster of similar data records corresponding to each representative data record in a selected set of representative data records to generate candidate values for the requested data field of the new data record, and generating a candidate score for each of the candidate values using the distance score and the confidence score to use in providing a recommended candidate value.


