Data Lake Auto-Tagging with Statistical Summaries and Exemplar Patterns
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Efficient and accurate tagging of large volumes of data, particularly custom data types, in data lakes is challenging due to the vast amount of data and non-standard formats, which complicates data governance and discovery.
Innovation Solution
A system that automatically generates a statistical summary of a data lake, interactively receives an exemplar set of data to determine a data-tagging pattern, and minimizes user interaction for efficient and accurate auto-tagging by filtering out under- and over-generalizing patterns.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual tagging methods are used for data in data lakes, then tagging accuracy can be maintained, but the time and labor required increases significantly due to the vast amount of data
Solution Approach 1:
The system performs preliminary analysis by generating a statistical summary of the data lake before actual tagging. This includes analyzing data distributions, identifying patterns, and pre-processing data to facilitate automated tagging, thereby reducing the time required for the main tagging operation while maintaining accuracy
Solution Approach 2:
The system introduces an intermediary statistical summary as a mediator between the raw data and the tagging process. This statistical summary captures essential characteristics of the data without requiring full manual inspection, enabling automated tagging to achieve accuracy comparable to manual methods while significantly reducing time investment
2Productivity
If automated tagging is implemented without user interaction, then processing speed increases, but tagging accuracy decreases due to inability to handle custom data types
Solution Approach 1:
The system applies partial automation by automatically generating statistical summaries and candidate tags, but requires minimal user interaction to review and confirm tags for custom data types. This partial approach maintains high productivity while ensuring accuracy through targeted human oversight where needed
Solution Approach 2:
The system incorporates feedback mechanisms where user interactions with automatically generated tags provide learning signals. The system uses this feedback to refine its statistical analysis and improve future automated tagging decisions, gradually increasing accuracy while maintaining efficiency
3Measurement precision
If comprehensive statistical analysis is performed on the entire data lake, then tagging accuracy improves, but computational resources and time required increase
Solution Approach 1:
The system segments the data lake analysis into manageable statistical summaries organized by data characteristics and patterns. Rather than analyzing all data uniformly, it divides the analysis into focused statistical components that can be processed efficiently while still capturing the essential patterns needed for accurate tagging
Solution Approach 2:
The system changes the parameters of analysis by focusing statistical computation on key data characteristics rather than exhaustive analysis of all data points. It transforms the problem from analyzing raw data volumes to analyzing statistical summaries with condensed information, reducing computational resource requirements while maintaining pattern recognition accuracy
4Measurement precision
If extensive user interaction is required for pattern selection, then tagging accuracy improves, but ease of operation deteriorates
Solution Approach 1:
The system applies local quality by providing detailed, focused interaction options only where needed for custom data types rather than requiring uniform user input across all data. For standard data types, tagging is fully automated, while custom types receive targeted user guidance, improving ease of operation while maintaining accuracy where it matters most
Data Source
AI summary
Systems and methods relate to auto-tagging of data in a data lake or a data storage. Generating a statistical summary of the data lake and interactively receiving data in a selected column of an exemplar data addresses an issue of efficiently and accurately auto-tagging data in a data lake. The present disclosure automatically generates a statistical summary of the data lake using a lightweight off-line processing. A graphical user interface interactively receives an exemplar data file with a selection of a column in the exemplar data file. A list of candidate data-tagging patterns is generated based on the statistical summary and updates the list by removing candidate data-tagging patterns that under-generalize the data. The present disclosure determines a data-tagging pattern by selecting a candidate data-tagging profile from the list based on having the least number of matching columns in the data lake.


