Unsupervised ML Master Data Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Maintaining accurate master data is challenging due to the presence of errors such as outdated values, inconsistencies, and typos, which are time-consuming and prone to errors when manually corrected.

Innovation Solution

The implementation of unsupervised master data correction using supervised machine learning, where machine learning models are applied to selected columns of a master data table to predict values, providing indications of recommended values, probabilities, and mismatches, facilitating both manual and automatic correction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual review and correction of master data is performed, then data quality can be maintained, but time consumption increases and errors remain prone

Engineering Contradiction:
Improvedata qualityVSAvoidtime consumption
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system enables self-service data correction by automatically detecting errors in master data and generating correction recommendations. The machine learning model analyzes data patterns, identifies anomalies, and suggests corrections without requiring manual intervention, allowing the system to correct its own data quality issues autonomously.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual review process with an automated machine learning-based system. The ML model processes data, detects errors, and generates corrections algorithmically, substituting human labor with an automated intelligent system that operates continuously without time loss.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual review and correction of master data is performed, then data accuracy can be maintained, but the process becomes error-prone

Engineering Contradiction:
Improvedata accuracyVSAvoidhuman error
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system substitutes manual correction processes with automated machine learning algorithms that objectively analyze data patterns and generate corrections. This eliminates human error by using consistent algorithmic logic rather than human judgment, which can be subjective and prone to mistakes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The machine learning model continuously learns from data patterns and provides feedback on potential errors. The system analyzes master data, identifies anomalies based on learned patterns, and generates corrections that can be reviewed or automatically applied, creating a feedback loop that improves data accuracy without human intervention.

Inventive Principle:
Principle #23Feedback

3Extent of automation

If unsupervised machine learning is used for data correction, then manual intervention is reduced, but model training complexity increases

Engineering Contradiction:
Improveautomation levelVSAvoidmodel training complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-processing the master data to prepare it for machine learning analysis. This includes data cleaning, feature engineering, and creating training datasets from historical data, which simplifies the subsequent model training process and enables automation without excessive complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies parameter changes by transforming the master data into appropriate formats and features that machine learning models can process effectively. This involves modifying data types, creating new features from existing data, and adjusting data representations to optimize model training and automation performance.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250139533A1Machine-learning-based unsupervised data correction
Publication Date: 2025.05.01 SAP SE
  • US20250139533A1 patent drawing
  • US20250139533A1 patent drawing
  • US20250139533A1 patent drawing

AI summary

Technologies are described for correcting data, such as master data, in an unsupervised manner using supervised machine learning. Correction of master data can involve receiving a table containing unlabeled master data. Machine learning models are applied to the fields of one or more columns of the table to predict values of the fields, and the machine learning models use unsupervised learning. For example, a machine learning model can be applied to a particular field of a particular column to predict the value of the particular field. The machine learning model uses the fields of other columns as features. Results of applying the machine learning models include indications of recommended values, indications of probabilities of the recommended values, and indications of which original values do not match their respective recommended values. The results can be used to perform manual and/or automatic correction of the master data.