Metadata Repository for Big Data Lifecycle Management
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current Big Data solutions, particularly those based on Hadoop, face challenges in efficiently managing and validating large datasets, lacking user-friendly applications for novice technologists and business users, and struggling with data validation and quantification, leading to inefficiencies in data loading, preparation, and analysis.
Innovation Solution
A data management platform with a metadata repository that tracks and manages all aspects of the data lifecycle, including storage, access controls, encryption, compression, and data lineage, providing self-service features through a GUI for analysts to find, select, and customize data for analysis, and automating data management processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If Hadoop-based Big Data solutions are used to process large datasets, then data processing capability is improved, but data validation and quantification become deficient
Solution Approach 1:
The patent introduces a metadata repository as an intermediary layer between the Hadoop data processing system and users. This metadata repository stores structured information about data sources, quality metrics, and validation rules, enabling reliable data validation and quantification without compromising the high processing capability of Hadoop. The metadata acts as a mediator that bridges the gap between raw data processing power and data quality assurance.
2Adaptability or versatility
If flexible schemas are used to handle unstructured and semi-structured data, then data versatility is improved, but data management complexity increases
Solution Approach 1:
The patent segments data management into two distinct layers: a flexible data storage layer that handles various data formats using Hadoop's flexible schemas, and a structured metadata management layer that organizes information about the data. This segmentation allows the system to maintain versatility in handling different data formats while reducing management complexity through the structured metadata layer that provides a consistent interface for data operations.
3Speed
If data is loaded without validation into HDFS, then data loading speed is improved, but data quality control deteriorates
Solution Approach 1:
The patent implements preliminary validation actions by storing metadata about data quality requirements, validation rules, and quality metrics in the metadata repository before data is loaded into HDFS. This preliminary action enables quality control to be prepared in advance, allowing fast data loading to proceed while maintaining quality standards through pre-defined validation criteria stored in the metadata layer.
4Productivity
If custom tools and technologies are integrated for data extraction and cleansing, then data preparation capability is improved, but administrative overhead increases
Solution Approach 1:
The patent creates a universal metadata repository that serves multiple functions: storing data source information, quality metrics, validation rules, and lineage information. This multi-functional metadata system replaces the need for multiple custom tools for data extraction, cleansing, and validation, thereby improving data preparation capability while reducing administrative overhead by consolidating these functions into a single unified metadata management infrastructure.
Data Source
AI summary
An analytical computing environment for large data sets comprises a software platform for data management. The platform provides various automation and self-service features to enable those users to rapidly provision and manage an agile analytics environment. The platform leverages a metadata repository, which tracks and manages all aspects of the data lifecycle. The repository maintains various types of platform metadata including, for example, status information (load dates, quality exceptions, access rights, etc.), definitions (business meaning, technical formats, etc.), lineage (data sources and processes creating a data set, etc.), and user data (user rights, access history, user comments, etc.). Within the platform, the metadata is integrated with all platform services, such as load processing, quality controls and system use. As the system is used, the metadata gets richer and more valuable, supporting additional automation and quality controls.


