Unsupervised ML Master Data Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Maintaining accurate master data is challenging due to the presence of errors such as outdated values, inconsistencies, and typos, which are time-consuming and prone to errors when manually corrected.
Innovation Solution
The implementation of unsupervised master data correction using supervised machine learning, where machine learning models are applied to selected columns of a master data table to predict values, providing indications of recommended values, probabilities, and mismatches, facilitating both manual and automatic correction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual review and correction of master data is performed, then data quality can be maintained, but time consumption increases and errors remain prone
Solution Approach 1:
The system enables self-service data correction by automatically detecting errors in master data and generating correction recommendations. The machine learning model analyzes data patterns, identifies anomalies, and suggests corrections without requiring manual intervention, allowing the system to correct its own data quality issues autonomously.
Solution Approach 2:
The patent replaces the mechanical manual review process with an automated machine learning-based system. The ML model processes data, detects errors, and generates corrections algorithmically, substituting human labor with an automated intelligent system that operates continuously without time loss.
2Reliability
If manual review and correction of master data is performed, then data accuracy can be maintained, but the process becomes error-prone
Solution Approach 1:
The system substitutes manual correction processes with automated machine learning algorithms that objectively analyze data patterns and generate corrections. This eliminates human error by using consistent algorithmic logic rather than human judgment, which can be subjective and prone to mistakes.
Solution Approach 2:
The machine learning model continuously learns from data patterns and provides feedback on potential errors. The system analyzes master data, identifies anomalies based on learned patterns, and generates corrections that can be reviewed or automatically applied, creating a feedback loop that improves data accuracy without human intervention.
3Extent of automation
If unsupervised machine learning is used for data correction, then manual intervention is reduced, but model training complexity increases
Solution Approach 1:
The system performs preliminary actions by pre-processing the master data to prepare it for machine learning analysis. This includes data cleaning, feature engineering, and creating training datasets from historical data, which simplifies the subsequent model training process and enables automation without excessive complexity.
Solution Approach 2:
The patent applies parameter changes by transforming the master data into appropriate formats and features that machine learning models can process effectively. This involves modifying data types, creating new features from existing data, and adjusting data representations to optimize model training and automation performance.
Data Source
AI summary
Technologies are described for correcting data, such as master data, in an unsupervised manner using supervised machine learning. Correction of master data can involve receiving a table containing unlabeled master data. Machine learning models are applied to the fields of one or more columns of the table to predict values of the fields, and the machine learning models use unsupervised learning. For example, a machine learning model can be applied to a particular field of a particular column to predict the value of the particular field. The machine learning model uses the fields of other columns as features. Results of applying the machine learning models include indications of recommended values, indications of probabilities of the recommended values, and indications of which original values do not match their respective recommended values. The results can be used to perform manual and/or automatic correction of the master data.


