Automated Data Quality Management via Self-Service Rule Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data quality management tools are inefficient, requiring manual processes, excessive human capital, and resources, and are not automated, leading to poor data quality assessment and decision-making in organizations, especially for large datasets, resulting in significant costs and reputational damage.

Innovation Solution

A system and method for automated data quality management that includes extracting data elements, generating rules based on data profiles, assessing data quality, detecting defects, and transmitting analysis for user display, utilizing a processor and memory with a user-friendly interface for streamlined data monitoring and analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual data quality management processes are used, then data quality can be assessed, but excessive human capital and time are required, reducing productivity

Engineering Contradiction:
Improvedata quality assessmentVSAvoidhuman capital efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables automated self-assessment of data quality through machine learning models that automatically evaluate data elements without requiring manual human intervention. The system extracts data elements, generates data profiles, creates assessment rules, and detects defects autonomously, allowing the data quality management process to serve itself rather than relying on human operators.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical processes with automated computational systems. Machine learning models and algorithms substitute human analysts in performing data quality assessment tasks, including extracting data elements, generating profiles, creating rules, and detecting defects, thereby eliminating the need for excessive human capital while maintaining or improving assessment precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated rule-based workflows are implemented, then productivity increases, but system complexity increases

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by automatically generating data profiles and creating assessment rules before actual data quality evaluation begins. Machine learning models pre-process the data to identify patterns and generate appropriate quality rules, which simplifies the subsequent assessment process and reduces the complexity of manual rule creation while maintaining high productivity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically adjusts parameters such as data quality thresholds, assessment criteria, and rule priorities based on the specific characteristics of the data being analyzed. This adaptability allows the system to handle different data types and quality requirements without requiring completely different system configurations, thereby managing complexity while maintaining high productivity across diverse scenarios.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If comprehensive data monitoring is performed, then data quality improves, but significant human capital and costs are required over long periods

Engineering Contradiction:
Improvedata qualityVSAvoidmonitoring time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system enables continuous automated monitoring of data quality without interruption or manual intervention. Machine learning models continuously evaluate data elements as they are processed, maintaining constant surveillance of data quality metrics. This continuous automated action ensures high data quality reliability without requiring sustained human capital investment over long periods.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system implements automated feedback loops where data quality assessment results are continuously fed back into the system to refine and improve future assessments. Machine learning models learn from detected defects and patterns, automatically adjusting assessment rules and priorities. This self-correcting feedback mechanism maintains high data quality reliability while eliminating the need for continuous human monitoring and analysis.

Inventive Principle:
Principle #23Feedback

4Ease of operation

If manual rule implementation is used, then ease of operation is maintained, but data quality improvement is limited

Engineering Contradiction:
Improvemanual operation simplicityVSAvoiddata quality
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system introduces machine learning models as intermediaries between simple manual operations and complex data quality assessment. Users can provide basic input or high-level guidelines, and the machine learning intermediary automatically generates comprehensive assessment rules, creates data profiles, and performs detailed quality evaluation. This intermediary approach maintains operational simplicity for users while achieving high data quality improvement that would otherwise require complex manual processes.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11030167B2Systems and methods for providing data quality management
Publication Date: 2021.06.08 CAPITAL ONE SERVICES LLC
  • US11030167B2 patent drawing
  • US11030167B2 patent drawing
  • US11030167B2 patent drawing

AI summary

A system for providing data quality management may include a processor configured to execute instructions to: extract a plurality of first data elements from a data source; generate a data profile based on the first data elements; automatically create a first set of rules based on the first data elements and the data profile, the first set of rules assessing data quality according to a threshold; generate a second set of rules based on the first data elements and the first set of rules; extract a plurality of second data elements; assess the second data elements based on a comparison of the second data elements to the second set of rules; detect defects based on the comparison; analyze data quality according to the detected defects; and transmit signals representing the data quality analysis to a client device for display to a user.