Active Entity Resolution Model for Procurement Data Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Duplicate data records in procurement and supply chain systems cause data messiness and incompleteness, making it difficult to search for suppliers and source goods or services efficiently.

Innovation Solution

An Active Entity Resolution (AER) model system that detects duplicate records, clusters them, generates canonical records, and uses machine learning to provide recommendations for new data records by matching them with master data, thereby improving data integrity and search functionality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If duplicate data records are allowed to exist in the system, then the quantity of data records increases, but data quality and search efficiency deteriorate

Engineering Contradiction:
Improvequantity of data recordsVSAvoiddata quality
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system extracts duplicate records from the data set using entity resolution algorithms. The AER model identifies and separates duplicate supplier records from unique records, allowing the system to maintain a complete data set while eliminating the harmful effects of duplication through selective extraction of duplicate instances.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system changes the state of data records by assigning confidence scores and similarity metrics to identify duplicates. By transforming the data through machine learning models that calculate match probabilities, the system can distinguish between unique and duplicate records without deleting any data, thus maintaining quantity while improving quality.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If manual processes are used to identify and manage duplicate records, then data accuracy can be maintained, but time consumption and operational complexity increase

Engineering Contradiction:
Improvedata accuracyVSAvoidtime consumption
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs self-service by automatically detecting, clustering, and managing duplicate records without human intervention. The AER model autonomously processes data records, identifies duplicates through machine learning, and generates recommendations, eliminating the need for manual data cleaning while maintaining high accuracy through algorithmic consistency.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical processes with an automated machine learning system. The AER model uses machine learning algorithms to substitute human analysts, automatically comparing records, calculating similarity scores, and identifying duplicates, thereby maintaining data accuracy while dramatically reducing time consumption.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Device complexity

If traditional entity resolution methods are used, then implementation simplicity is maintained, but resolution accuracy and automation capability deteriorate

Engineering Contradiction:
Improvesystem simplicityVSAvoidautomation capability
Core Design Contradiction:
Device complexityVSExtent of automation

Solution Approach 1:

The AER model serves multiple functions within a single unified system: it performs entity resolution, duplicate detection, data clustering, and recommendation generation. This multi-functional approach maintains relative system simplicity while achieving high automation capability, as the single model handles various data quality tasks that would otherwise require multiple separate tools.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11720601B2Active entity resolution model recommendation system
Publication Date: 2023.08.08 SAP SE
  • US11720601B2 patent drawing
  • US11720601B2 patent drawing
  • US11720601B2 patent drawing

AI summary

Systems and methods are provided for accessing master data comprising a plurality of representative data records, where each representative data record represents a cluster of similar data records, and each similar data record has a confidence score indicating a confidence level that the similar data record corresponds to the cluster, and comparing a new data record to each representative data record of the plurality of representative data records using a machine learning model to generate a distance score. The systems and methods further provide for analyzing the cluster of similar data records corresponding to each representative data record in a selected set of representative data records to generate candidate values for the requested data field of the new data record, and generating a candidate score for each of the candidate values using the distance score and the confidence score to use in providing a recommended candidate value.