Data Management Platform Automating Metadata Validation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data analytics professionals face challenges in managing disparate data sources with different formats and formatting requirements, leading to inefficient use of computing resources and network resources due to manual updates, metadata validation, and job request failures.
Innovation Solution
A data management platform that configures a data environment to meet application requirements, performs metadata validation, and uses machine learning to monitor job requests, thereby conserving computing resources and automating updates and metadata schema management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual updates and metadata validation are performed for disparate data sources, then data integrity can be maintained, but computing resources and network resources are consumed inefficiently
Solution Approach 1:
The system performs preliminary actions by configuring the data environment with metadata schemas and validation rules before data is loaded or processed. This includes pre-defining metadata schemas for different data sources and applications, and setting up validation logic in advance, so that automated validation occurs during data operations rather than requiring manual intervention, thereby maintaining data integrity while reducing resource consumption.
Solution Approach 2:
The system enables self-service through automated metadata validation and data reconciliation processes. The configured data environment automatically validates metadata against defined schemas, performs data reconciliation between sources, and generates validation logs without manual intervention. This automation maintains data integrity while significantly reducing the computing and network resources that would be consumed by manual processes.
2Manufacturing precision
If manual metadata validation is performed for each data source, then data format compliance can be ensured, but time and computing resources are wasted
Solution Approach 1:
The system implements universality by creating a configurable data environment that uses a master metadata schema that can serve multiple data sources and applications. The metadata validation framework is designed to be reusable across different data sources, allowing the same validation logic to be applied universally rather than requiring separate manual validation for each source, thus ensuring data format compliance while reducing validation time.
Solution Approach 2:
The system incorporates feedback mechanisms through automated metadata validation that provides immediate results. The validation process compares metadata against configured schemas and generates validation logs with pass/fail indicators, enabling rapid identification and correction of format compliance issues without time-consuming manual review, thereby ensuring precision while minimizing time loss.
3Measurement precision
If data environment is configured manually to meet application requirements, then data analytics accuracy can be improved, but device complexity and operational overhead increase
Solution Approach 1:
The system applies dynamics by making the data environment configuration flexible and adaptable rather than static and rigid. The metadata schemas and validation rules can be dynamically configured and updated to meet different application requirements without requiring complex manual reconfiguration. The system can adapt to changing data sources and application needs while maintaining data analytics accuracy through automated validation.
Solution Approach 2:
The system introduces an intermediary layer in the form of a configurable data environment that sits between raw data sources and analytics applications. This intermediary automatically handles metadata validation, data reconciliation, and format standardization, thereby improving data analytics accuracy while shielding users from the complexity of underlying configuration details through automation.
4Stability of the object's composition
If automated data reconciliation is implemented across multiple data sources, then data consistency can be maintained, but system complexity increases
Solution Approach 1:
The system applies segmentation by dividing the data environment into modular components, including separate metadata schemas for different data sources, individual validation rules for each source, and organized validation logs. This segmentation allows automated data reconciliation to be implemented in a structured, manageable way across multiple data sources while maintaining data consistency, as each segment can be configured and validated independently.
Data Source
AI summary
A data management platform may receive an environment configuration for a data environment to be implemented in a data structure, wherein the environment configuration includes requirements of an application. The data management platform may configure, based on the environment configuration, the data environment, to generate a configured data environment. The data management platform may deploy the configured data environment in the data structure. The data management platform may perform one or more tests on data stored in the configured data environment in the data structure to generate one or more test results based on performing the one or more tests on the data. The data management platform may update, based on the one or more test results, the configured data environment, to generate an updated configured data environment, wherein the updated configured data environment meets the requirements of the application.


