ML Data Asset Profiling Across Distributed Hadoop Clusters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Enterprises face challenges in managing vast amounts of data spread across multiple clusters and repositories due to the lack of a uniform method for identifying user behavior, data usage density, sensitivity, and quality, particularly in hybrid storage systems without a schema, making it difficult to classify and govern data assets effectively.

Innovation Solution

A data management system utilizing advanced machine learning algorithms to automatically locate, identify, and categorize data assets across Hadoop multi-cluster environments, integrating with tools like Apache Atlas and Apache Ranger for metadata management and security, providing a unified interface for data stewards to manage and govern distributed data assets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual data classification methods are used by data stewards, then data assets can be managed with some level of control, but the process becomes arduous, time-consuming, and lacks complete coverage

Engineering Contradiction:
Improvedata classification efficiencyVSAvoidtime required for data management tasks
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system enables automated self-service data classification through machine learning algorithms that automatically profile, classify, and tag data assets without requiring manual intervention from data stewards. The ML models autonomously analyze data patterns, identify sensitive information, and apply appropriate classification labels, freeing data stewards from manual classification tasks while maintaining comprehensive coverage.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical processes (data stewards manually examining and classifying data) with automated machine learning systems. The ML algorithms automatically perform data profiling, sensitivity detection, and classification tasks that previously required human effort, significantly improving efficiency and reducing time loss.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Extent of automation

If automated machine learning classification is implemented, then data classification efficiency and coverage improve, but system complexity increases

Engineering Contradiction:
Improveautomated data classificationVSAvoidsystem architecture complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The system introduces an intermediary layer of machine learning models and profiling services that sit between the raw data assets and the data management system. These ML intermediaries handle the complex classification logic, automatically generating metadata and classification labels that simplify the overall data management process. The intermediary ML layer absorbs the complexity, presenting a simplified interface to data stewards.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If data is stored across distributed file systems without uniform identification methods, then data storage flexibility is maintained, but data governance and security management become significantly challenging

Engineering Contradiction:
Improvedata storage flexibilityVSAvoiddata governance and security management
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent implements a universal machine learning-based classification system that works across diverse distributed file systems and data formats. The ML models are designed to handle multiple data types, storage formats, and file systems uniformly, applying consistent classification and tagging rules across the entire distributed environment. This universal approach maintains storage flexibility while enabling centralized governance and security management.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Loss of information

If comprehensive data profiling and classification are performed across all data assets, then complete data visibility and governance are achieved, but computational resources and processing time increase

Engineering Contradiction:
Improvedata visibility and awarenessVSAvoidcomputational resource consumption
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The system implements selective data profiling that focuses computational resources on the most critical data assets and classification categories. Rather than uniformly profiling every single data asset with equal depth, the ML system prioritizes profiling of sensitive data, frequently accessed data, and data in critical storage locations. This partial action approach achieves sufficient data visibility for effective governance while reducing overall computational resource consumption.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10983963B1Automated discovery, profiling, and management of data assets across distributed file systems through machine learning
Publication Date: 2021.04.20 HORTONWORKS INC
  • US10983963B1 patent drawing
  • US10983963B1 patent drawing
  • US10983963B1 patent drawing

AI summary

Embodiments for locating, identifying and categorizing data-assets through advanced machine learning algorithms implemented by profiler components across Hadoop and Hadoop Compatible File Systems, databases and in-memory objects automatically and periodically to provide a visual representation of the category of data infrastructure distributed across data-centers and multiple clusters, for the purposes of enriching data quality, enabling data discovery and improving outcomes from downstream systems.