Data Classification Ontology for Data Lake Governance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data lake management systems face challenges in efficiently classifying and governing diverse data types, leading to data swamps rather than effective data lakes due to deficiencies in data governance, particularly in accurately identifying data and mapping it to enterprise-wide business standards.

Innovation Solution

A system that uses natural language understanding processing to extract training data elements from structured text, identify associations, and build a logical relationship data classification knowledge base ontology, connecting business classes to extracted data elements, enabling automated data classification and governance across various data sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual data classification methods are used, then data governance accuracy can be maintained, but classification efficiency and productivity are poor

Engineering Contradiction:
Improvedata classification efficiencyVSAvoidmanual classification effort
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The system enables self-service data classification by automatically extracting data elements, identifying associations, and building classification ontologies without requiring manual intervention. The natural language understanding processing autonomously performs classification tasks that would otherwise require human analysts, directly resolving the contradiction between productivity and automation extent.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical classification processes with automated natural language understanding processing. The system uses computational methods to extract and classify data elements, substituting human cognitive processes with automated algorithms that achieve both high productivity and maintained accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated data classification is implemented, then classification efficiency improves, but data mapping accuracy to business standards may deteriorate

Engineering Contradiction:
Improveclassification efficiencyVSAvoiddata mapping accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system incorporates feedback mechanisms where extracted data elements are validated against business standards and taxonomy classifications. The natural language understanding processing continuously refines its mappings based on feedback from business terminology, ensuring that automated classification maintains high accuracy while achieving efficiency improvements.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs preliminary extraction and identification of data elements before final classification. By pre-processing the data to extract training set data elements and identify associations with business classes, the system prepares accurate mappings that maintain precision while enabling efficient automated classification in subsequent steps.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If comprehensive data governance is implemented, then data quality and accuracy improve, but system complexity increases

Engineering Contradiction:
Improvedata governance qualityVSAvoidgovernance system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system extracts and isolates specific data elements and their associations from the broader data corpus. By extracting training set data elements and creating discrete question-business class associations, the system manages complexity through modular extraction rather than attempting to process entire datasets monolithically, thereby maintaining governance quality while reducing operational complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the data governance process into distinct stages: extracting data elements, identifying associations, building ontologies, and classifying data. This segmentation allows comprehensive governance to be achieved through manageable, discrete operations rather than a single complex system, reducing overall system complexity while maintaining reliability.

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If natural language understanding processing is used, then extraction accuracy of data elements improves, but processing time increases

Engineering Contradiction:
Improvedata element extraction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system applies partial natural language understanding processing targeted at specific data elements and business classes rather than processing entire datasets uniformly. By focusing the computationally intensive NLP operations only where needed and using pre-extracted training data where possible, the system maintains high extraction accuracy while reducing overall processing time through selective application of complex analysis.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11893500B2Data classification for data lake catalog
Publication Date: 2024.02.06 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11893500B2 patent drawing
  • US11893500B2 patent drawing
  • US11893500B2 patent drawing

AI summary

Aspects include processors configured to (or include program code that causes a processor to) provide for data classifier devices that extract from structured text business data inputs, via natural language understanding processing, training set data elements (for example, training keywords, training concepts, training entities, and/or training taxonomy classifications, etc.). The aspects identify associations within the structured training business data of each of a plurality of business class categories with respective ones of the extracted training set data elements; and build a logical relationship data classification training knowledge base ontology that connects ones of the business classes to respective associated ones of the extracted training data elements as questions, into a plurality of knowledge base ontology question-business class associations.