Data Classification Ontology for Data Lake Governance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data lake management systems face challenges in efficiently classifying and governing diverse data types, leading to data swamps rather than effective data lakes due to deficiencies in data governance, particularly in accurately identifying data and mapping it to enterprise-wide business standards.
Innovation Solution
A system that uses natural language understanding processing to extract training data elements from structured text, identify associations, and build a logical relationship data classification knowledge base ontology, connecting business classes to extracted data elements, enabling automated data classification and governance across various data sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual data classification methods are used, then data governance accuracy can be maintained, but classification efficiency and productivity are poor
Solution Approach 1:
The system enables self-service data classification by automatically extracting data elements, identifying associations, and building classification ontologies without requiring manual intervention. The natural language understanding processing autonomously performs classification tasks that would otherwise require human analysts, directly resolving the contradiction between productivity and automation extent.
Solution Approach 2:
The patent replaces manual mechanical classification processes with automated natural language understanding processing. The system uses computational methods to extract and classify data elements, substituting human cognitive processes with automated algorithms that achieve both high productivity and maintained accuracy.
2Productivity
If automated data classification is implemented, then classification efficiency improves, but data mapping accuracy to business standards may deteriorate
Solution Approach 1:
The system incorporates feedback mechanisms where extracted data elements are validated against business standards and taxonomy classifications. The natural language understanding processing continuously refines its mappings based on feedback from business terminology, ensuring that automated classification maintains high accuracy while achieving efficiency improvements.
Solution Approach 2:
The patent performs preliminary extraction and identification of data elements before final classification. By pre-processing the data to extract training set data elements and identify associations with business classes, the system prepares accurate mappings that maintain precision while enabling efficient automated classification in subsequent steps.
3Reliability
If comprehensive data governance is implemented, then data quality and accuracy improve, but system complexity increases
Solution Approach 1:
The system extracts and isolates specific data elements and their associations from the broader data corpus. By extracting training set data elements and creating discrete question-business class associations, the system manages complexity through modular extraction rather than attempting to process entire datasets monolithically, thereby maintaining governance quality while reducing operational complexity.
Solution Approach 2:
The patent segments the data governance process into distinct stages: extracting data elements, identifying associations, building ontologies, and classifying data. This segmentation allows comprehensive governance to be achieved through manageable, discrete operations rather than a single complex system, reducing overall system complexity while maintaining reliability.
4Measurement precision
If natural language understanding processing is used, then extraction accuracy of data elements improves, but processing time increases
Solution Approach 1:
The system applies partial natural language understanding processing targeted at specific data elements and business classes rather than processing entire datasets uniformly. By focusing the computationally intensive NLP operations only where needed and using pre-extracted training data where possible, the system maintains high extraction accuracy while reducing overall processing time through selective application of complex analysis.
Data Source
AI summary
Aspects include processors configured to (or include program code that causes a processor to) provide for data classifier devices that extract from structured text business data inputs, via natural language understanding processing, training set data elements (for example, training keywords, training concepts, training entities, and/or training taxonomy classifications, etc.). The aspects identify associations within the structured training business data of each of a plurality of business class categories with respective ones of the extracted training set data elements; and build a logical relationship data classification training knowledge base ontology that connects ones of the business classes to respective associated ones of the extracted training data elements as questions, into a plurality of knowledge base ontology question-business class associations.


