Automated Data Quality Task Ranking and Routing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing volume of data in data lakes makes manual inspection and cleansing impossible, leading to a high number of data quality exceptions that require human intervention, while not all data sets or problems have equal importance, and stewards' time is limited.
Innovation Solution
A system that automatically ranks and routes data quality remediation tasks by computing a score based on the importance and cost of resolution, using data lineage and business classifications, and trained machine-learning models to predict steward capacity, optimizing steward workload.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual inspection and cleansing of data is performed, then data quality can be ensured, but the process becomes impossible as data volume increases
Solution Approach 1:
The system enables data quality management to serve itself through automated scoring and routing mechanisms. The data quality system automatically evaluates datasets, computes priority scores, and routes tasks without requiring manual intervention at each step, allowing the system to scale with data volume while maintaining quality standards
Solution Approach 2:
The system transforms the data quality management process by introducing computed priority scores as a new parameter. This score, derived from multiple factors including data lineage and business classification, enables automated ranking and routing of data quality tasks, replacing manual prioritization and enabling scalable processing
2Reliability
If all data quality problems are treated equally, then comprehensive coverage is achieved, but steward time is wasted on low-priority issues
Solution Approach 1:
The system applies differentiated quality management to different data quality problems based on their local characteristics. By computing priority scores that reflect the specific importance, risk, and business context of each dataset and problem type, the system enables stewards to focus on high-priority issues while maintaining appropriate coverage across all data quality concerns
Solution Approach 2:
The system performs preliminary evaluation and scoring of all data quality problems before steward intervention. By pre-computing priority scores based on data lineage, business classification, and problem characteristics, the system prepares and ranks tasks in advance, enabling stewards to immediately focus on the most critical issues without manual assessment
3Productivity
If automated ranking and routing is implemented, then steward efficiency is improved, but system complexity increases
Solution Approach 1:
The system employs a universal scoring framework that handles multiple data quality problem types, datasets, and routing scenarios through a single priority score computation mechanism. This multi-functional approach consolidates what would otherwise require multiple specialized systems into one unified platform, managing complexity while maintaining high steward efficiency
Data Source
AI summary
In an approach for automatically ranking and routing data quality remediation tasks, a processor analyzes a data set ingested by a repository to produce a set of data quality problems. A processor computes a score for each data quality problem of the set of data quality problems. A processor identifies a route to send each data quality problem of the set of data quality problems. A processor exports each data quality problem according to the score and the route.


