Data Quality Analysis System for Big Data Metrics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current distributed cluster computing systems focused on capturing data lack the ability to effectively analyze and evaluate the quality of big data sets, including determining what data to store, how data links together, and establishing appropriate data models and standards.
Innovation Solution
A method and apparatus that retrieve rules from an asset catalog, generate metrics based on counter information from a data set, and evaluate these metrics using a processor to identify and notify on quality issues, with features like ETL operations, custom collectors, and historical metric analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If distributed cluster computing systems are used to manage big data sets, then data storage capacity and processing power are improved, but the ability to analyze and evaluate data quality deteriorates
Solution Approach 1:
The patent introduces data quality metrics as an intermediary layer between the distributed storage system and the data itself. These metrics (completeness, conformity, validity, timeliness, uniqueness) serve as mediators that capture quality information without requiring complex analysis of the entire data set, thus enabling quality evaluation in large-scale distributed systems.
Solution Approach 2:
The patent replaces traditional mechanical data quality analysis methods with automated metric-based evaluation. Instead of manually analyzing data quality through complex processing, the system uses predefined metrics and automated calculations to assess quality attributes, making the process scalable to big data environments.
2Ease of operation
If traditional data management tools are used, then data processing simplicity is maintained, but the ability to handle and analyze large data sets deteriorates
Solution Approach 1:
The patent segments data quality evaluation into distinct, manageable metrics (completeness, conformity, validity, timeliness, uniqueness). Each metric can be independently calculated and monitored, allowing simple operational procedures even for large data sets while maintaining comprehensive quality assessment.
3Measurement precision
If data quality analysis capabilities are added to distributed computing systems, then data quality evaluation is improved, but system complexity increases
Solution Approach 1:
The patent creates a universal data quality metrics framework that can be applied across multiple data types, storage systems, and analysis tools. This multi-functional approach allows the same metric infrastructure to serve diverse data quality evaluation needs, reducing overall system complexity through standardization.
Data Source
AI summary
A system is disclosed to evaluate data quality in a big data environment. An example method performed by the system includes retrieving one or more rules from an asset catalog. The method further includes retrieving, based on the one or more rules, counter information from a data set, and generating, by a processor, one or more metrics based on the one or more rules and the counter information. In addition, the method includes evaluating, by the processor, the one or more metrics based on the one or more rules. In an instance in which evaluation of a particular metric of the one or more metrics identifies an attribute value that exceeds a predetermined threshold, the method includes causing a notification message regarding the particular metric to be output. A corresponding apparatus and computer program product are also provided.


