Data Lake Metadata Classification With Guardrail Consistency Checks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data lake catalogs suffer from incomplete, inaccurate, and inefficient metadata assignment, leading to issues like overclassification, inconsistent naming conventions, and increased system resource usage due to reliance on third-party systems, which can result in security and integrity risks.
Innovation Solution
A method and system using classification models and data guardrails to automatically assign metadata values, including global classifiers, confidentiality sub-class classifiers, and sensitivity classifiers, with acronym expansion and description enrichment to ensure accuracy and consistency, reducing reliance on third-party systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual classification is used to assign metadata values to data elements, then flexibility and accuracy in handling complex cases can be maintained, but the process becomes time consuming and labor intensive
Solution Approach 1:
The system enables self-service automated metadata classification through machine learning models that automatically analyze data elements and assign appropriate metadata values without requiring manual intervention for every data element, thereby reducing time consumption while maintaining accuracy through continuous learning and validation mechanisms
Solution Approach 2:
The patent replaces the mechanical manual classification process with an automated electronic system using machine learning models, natural language processing, and algorithmic validation to assign metadata values, eliminating the need for manual human labor while improving consistency and speed of classification
2Reliability
If third-party systems are used for metadata classification, then external expertise and validation can be leveraged, but system resource usage increases and integration complexity arises
Solution Approach 1:
The patent merges multiple classification functions and validation mechanisms into a single integrated automated system that combines machine learning models, rule-based validation, and metadata management capabilities, eliminating the need for separate third-party systems while maintaining reliability through internal consistency checks
Solution Approach 2:
The system achieves multi-functionality by implementing a unified metadata classification platform that can handle various data types, classification schemes, and validation requirements within a single system, replacing multiple specialized third-party tools and reducing overall system resource consumption
3Reliability
If overclassification is applied to ensure data security, then security risks are mitigated, but data discoverability and access efficiency are reduced
Solution Approach 1:
The system applies local quality by assigning metadata classifications at the appropriate granularity level for each specific data element rather than applying uniform overclassification to entire datasets, allowing security measures to be precisely targeted where needed while maintaining ease of access for non-sensitive data
Solution Approach 2:
The patent dynamically adjusts classification parameters and metadata assignments based on the actual sensitivity and security requirements of each data element, using automated analysis to determine the minimum necessary classification level rather than applying fixed overclassification rules, thereby balancing security with discoverability
Data Source
AI summary
Various methods and processes, apparatuses or systems, and media for automatically assigning metadata values to data elements within a data lake catalog in an accurate and efficient manner are disclosed. The method includes: receiving a first data set that includes a plurality of data elements; using a first classification model to assign, to each data element, respective first metadata that includes a respective global classifier, a respective confidentiality sub-class classifier and a sensitivity classifier; determining, for each data element based on the corresponding global classifier and the corresponding confidentiality sub-class classifier, a respective confidence threshold; determining, for each data element based on the corresponding confidence threshold and the corresponding sensitivity classifier, a respective consistency value; and applying, to each data element, at least one data guardrail to check whether the data element is consistent with the assigned metadata.


