Zone-Based Data Governance with Automated Dataset Movement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data governance systems lack efficient and automated processes for managing and governing digital assets, leading to unnecessary time and resource expenditure in understanding data interconnectedness without proper visual or structural representation.
Innovation Solution
A zone-based database management system that creates and manages dataset zones, including transient, raw, trusted, refined, and analytical workspace zones, with data pipelines and machine learning for automated data movement and classification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional data governance systems are used, then data can be managed with basic policies and procedures, but the process is time-consuming and lacks automated efficiency
Solution Approach 1:
The system enables automated data governance through machine learning algorithms that autonomously classify datasets, determine zone assignments, and enforce policies without requiring manual intervention. The data governance system self-manages the complex tasks of data interconnectedness analysis and policy enforcement, eliminating time-consuming manual processes.
Solution Approach 2:
Traditional manual data governance processes are replaced with automated machine learning mechanisms. The system uses AI algorithms to substitute human analysis and decision-making in data classification, zone assignment, and policy enforcement, thereby increasing efficiency and reducing time expenditure.
2Ease of operation
If manual data classification and governance processes are used, then flexibility in handling different data types is maintained, but the process lacks structure and visual representation making it complex
Solution Approach 1:
The system segments data into distinct zones (e.g., raw data zone, cleaned data zone, analytics zone) with specific governance policies applied to each. This segmentation provides a structured visual representation of data flow and interconnectedness, making the complex governance process easier to understand and operate.
Solution Approach 2:
The system uses visual indicators and color-coding to represent different data zones, data types, and governance statuses. This visual representation transforms complex data interconnectedness into intuitive graphical displays, enhancing ease of operation while maintaining structural organization.
3Reliability
If basic data storage without zones is used, then system simplicity is maintained, but data security and compliance governance are insufficient
Solution Approach 1:
The system divides data storage into multiple secured zones, each with specific access controls and governance policies. This segmentation enhances data security and compliance by isolating sensitive data and applying targeted protection measures, while the modular zone structure manages complexity through standardized templates.
Solution Approach 2:
The system dynamically adjusts security parameters and access controls based on data characteristics, sensitivity levels, and governance policies. By changing security parameters automatically based on data zone assignments, the system achieves robust security and compliance without manual complexity.
4Productivity
If automated machine learning classification is implemented, then data governance speed and scale are improved, but the system requires training data and computational resources
Solution Approach 1:
The system performs preliminary training of machine learning models using historical data and predefined governance policies before actual data governance tasks are executed. This preliminary action enables the system to achieve high governance speed and scale for new data without requiring extensive real-time computational resources, as the models are pre-trained and ready for inference.
Data Source
AI summary
Systems and methods are disclosed for moving data based on data governance policies, wherein a plurality of datasets from a plurality of sources are received and stored into a data catalog. Predefined zones are generated, each having predefined policies. At least one common characteristic is determined for a first dataset and a second dataset. The system receives a request from an authorized user to move the first and second datasets into a particular zone and moves the datasets accordingly. The system then displays, via a graphical user interface, a representation depicting the first and second datasets. The predefined zones include a transient zone, a raw zone, a trusted zone, and a refined zone. The plurality of datasets move through the zones through a data pipeline that performs a data quality check to ensure that the data is moved through the zones according to the predefined polices.


