Data Quality Processing with ML Error Correction Feedback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Organizations face inefficiencies due to poor data quality, including corruption, inconsistency, and duplication, which lead to errors and increased processing resources when data is repeatedly fixed upon use, reducing the time data can be effectively utilized.
Innovation Solution
A data quality system that processes data from various sources, uses pre-processing techniques to prepare data, and applies machine learning to identify and correct errors, updating the source data to prevent repetitive fixing and enhance data usability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data is processed repeatedly to fix errors upon use, then data quality is improved, but processing resources are consumed and time is lost
Solution Approach 1:
The system performs preliminary data quality assessment and error correction before data is used by applications. A data quality system receives data from sources, assesses its quality using trained machine learning models, and corrects identified errors in advance, eliminating the need for repeated fixing when data is accessed later.
Solution Approach 2:
The system implements a feedback mechanism where data quality assessment results are used to trigger corrective actions. The machine learning model continuously learns from assessed data patterns and improves its ability to identify and correct errors, creating a closed-loop system that enhances data quality over time without increasing processing overhead.
2Productivity
If data quality assessment and correction is performed in advance, then processing resources are conserved, but system complexity increases
Solution Approach 1:
The patent introduces a data quality system as an intermediary component between data sources and applications. This mediator receives data from sources, performs quality assessment using machine learning models, corrects errors, and provides cleaned data to applications, thereby managing complexity in a modular fashion without requiring changes to existing data sources or applications.
Solution Approach 2:
The system employs self-service mechanisms where the machine learning model automatically assesses and corrects data quality issues without requiring manual intervention. The model is trained on historical data patterns and autonomously identifies and corrects errors, reducing the need for complex manual data management processes.
Data Source
AI summary
A first device may receive data from a set of second devices to be processed to determine a quality of the data. The data may include first data stored by the set of second devices, second data provided toward a third device, or third data related to fourth data. The first device may process the data using a first set of techniques to prepare the data for processing. The first device may process the data using a second set of techniques to improve the quality of the data and to form processed data. The first device may provide the processed data toward the set of second devices to replace the data stored by the set of second devices to permit the set of second devices to use the processed data. The first device may perform an action after providing the processed data toward the set of second devices.


