Interactive Attribute Trees for Data Quality Before Ingestion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data management techniques fail to consider data consumption and utility metrics, leading to inefficient data discovery and management in large datasets, particularly in data lakes, and lack integration of user and usage data for improved collaboration and decision-making.
Innovation Solution
Implement data quality, consumption, and utility metrics at attribute-, record-, and schema-level, leveraging interaction data to compute and update metrics for improved data ingestion, selection, and monitoring, using tools that facilitate navigation and discovery.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is ingested into a data lake without careful preparation, then data ingestion speed and volume increase, but data quality and usability deteriorate
Solution Approach 1:
The system performs preliminary data quality assessment by computing metrics on sample data before full ingestion, allowing early detection of quality issues and preventing ingestion of poor-quality data at scale
Solution Approach 2:
The system continuously monitors data quality metrics after ingestion and provides feedback to guide data cleaning, transformation, and improvement actions, creating a closed-loop data management process
2Loss of information
If data analysts explore large datasets manually, then they can discover insights, but the time and effort required increase significantly
Solution Approach 1:
The system replaces manual mechanical exploration with automated computational analysis by computing and visualizing data quality and consumption metrics across the dataset, enabling rapid assessment without manual inspection
Solution Approach 2:
The system introduces data quality metrics and consumption metrics as intermediary representations that mediate between the raw data and the analyst, providing summarized information about data characteristics without requiring direct examination of all data points
3Measurement precision
If comprehensive data quality metrics are computed for all data, then data quality assessment improves, but computational resources and time increase
Solution Approach 1:
The system computes data quality metrics on a sample of data (partial action) rather than all data, providing sufficient assessment information without the full computational cost of analyzing every record
Solution Approach 2:
The system segments the data assessment process by computing metrics at different levels (attribute-level, record-level, dataset-level) and only computing detailed metrics where needed based on higher-level summaries
Data Source
AI summary
Embodiments provide systems, methods, and computer storage media for management, assessment, navigation, and/or discovery of data based on data quality, consumption, and/or utility metrics. Data may be assessed using attribute-level and/or record-level metrics that quantify data: “quality”—the condition of data (e.g., presence of incorrect or incomplete values), its “consumption”—the tracked usage of data in downstream applications (e.g., utilization of attributes in dashboard widgets or customer segmentation rules), and/or its “utility”−a quantifiable impact resulting from the consumption of data (e.g., revenue or number of visits resulting from marketing campaigns that use particular datasets, storage costs of data). This data assessment may be performed at different stages of a data intake, preparation, and/or modeling lifecycle. For example, an interactive tree view may visually represent a nested attribute schema and attribute quality or consumption metrics to facilitate discovery of bad data before ingesting into a data lake.


