Data Quality Processing with ML Error Correction Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Organizations face inefficiencies due to poor data quality, including corruption, inconsistency, and duplication, which lead to errors and increased processing resources when data is repeatedly fixed upon use, reducing the time data can be effectively utilized.

Innovation Solution

A data quality system that processes data from various sources, uses pre-processing techniques to prepare data, and applies machine learning to identify and correct errors, updating the source data to prevent repetitive fixing and enhance data usability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is processed repeatedly to fix errors upon use, then data quality is improved, but processing resources are consumed and time is lost

Engineering Contradiction:
Improvedata qualityVSAvoidprocessing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary data quality assessment and error correction before data is used by applications. A data quality system receives data from sources, assesses its quality using trained machine learning models, and corrects identified errors in advance, eliminating the need for repeated fixing when data is accessed later.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements a feedback mechanism where data quality assessment results are used to trigger corrective actions. The machine learning model continuously learns from assessed data patterns and improves its ability to identify and correct errors, creating a closed-loop system that enhances data quality over time without increasing processing overhead.

Inventive Principle:
Principle #23Feedback

2Productivity

If data quality assessment and correction is performed in advance, then processing resources are conserved, but system complexity increases

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces a data quality system as an intermediary component between data sources and applications. This mediator receives data from sources, performs quality assessment using machine learning models, corrects errors, and provides cleaned data to applications, thereby managing complexity in a modular fashion without requiring changes to existing data sources or applications.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system employs self-service mechanisms where the machine learning model automatically assesses and corrects data quality issues without requiring manual intervention. The model is trained on historical data patterns and autonomously identifies and corrects errors, reducing the need for complex manual data management processes.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10379920B2Processing data to improve a quality of the data
Publication Date: 2019.08.13 ACCENTURE GLOBAL SOLUTIONS LTD
  • US10379920B2 patent drawing
  • US10379920B2 patent drawing
  • US10379920B2 patent drawing

AI summary

A first device may receive data from a set of second devices to be processed to determine a quality of the data. The data may include first data stored by the set of second devices, second data provided toward a third device, or third data related to fourth data. The first device may process the data using a first set of techniques to prepare the data for processing. The first device may process the data using a second set of techniques to improve the quality of the data and to form processed data. The first device may provide the processed data toward the set of second devices to replace the data stored by the set of second devices to permit the set of second devices to use the processed data. The first device may perform an action after providing the processed data toward the set of second devices.