Dataset Management System for Distributed Storage Placement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data management systems struggle to efficiently manage the lifecycle of large and dynamic datasets, particularly in scenarios where data is generated from multiple sources and locations, leading to increased complexity in monitoring and accessing data.
Innovation Solution
Implementing a dataset management system that monitors data ownership and use across an organization, allowing for efficient placement of data in appropriate storage locations and assignment of attributes for optimal use, by utilizing metadata and data queries to create and manage datasets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If data is stored in multiple geographic locations to improve accessibility and reduce latency, then data access speed is improved, but monitoring and managing data placement becomes overly cumbersome and complex
Solution Approach 1:
The patent introduces a data placement monitor as an intermediary component that automatically tracks data placement across multiple storage locations and generates placement reports. This mediator handles the complexity of monitoring distributed data, allowing the system to maintain fast local access while automatically managing the oversight of data across geographic locations without requiring manual intervention.
Solution Approach 2:
The system implements feedback mechanisms where the data placement monitor continuously observes data placement status and generates reports that feed back into the data management system. This feedback loop enables automatic adjustment and optimization of data placement strategies, reducing the manual monitoring burden while maintaining optimal data accessibility across distributed locations.
2Manufacturing precision
If manual data management is used to maintain data lifecycle policies, then data management accuracy is improved, but scalability is severely limited as data volumes grow to exabyte scales
Solution Approach 1:
The patent implements self-service mechanisms where the data management system automatically executes lifecycle policies based on monitored data characteristics and placement information. The system autonomously performs data tiering, archiving, and deletion operations according to predefined policies without requiring continuous manual intervention, thereby maintaining management accuracy while achieving scalability to exabyte-level data volumes.
Solution Approach 2:
The system employs preliminary action by pre-defining data lifecycle policies and placement rules before data growth occurs. These predetermined policies are automatically applied as data is ingested and placed, enabling the system to handle massive data volumes without requiring proportional increases in manual management resources, thus achieving both accuracy and scalability.
3Productivity
If data is placed locally to minimize latency for specific projects, then data access efficiency is improved, but optimizing storage placement across distributed systems becomes cumbersome
Solution Approach 1:
The data placement monitor acts as an intermediary that automatically analyzes data access patterns and placement requirements, generating optimized placement recommendations. This mediator handles the complexity of determining optimal local storage locations for different projects, allowing the system to achieve fast local data access without manual optimization efforts.
Solution Approach 2:
The patent replaces manual mechanical optimization processes with automated electronic monitoring and analysis systems. The data placement monitor uses automated algorithms to analyze access patterns and determine optimal storage locations, substituting manual trial-and-error placement methods with systematic, data-driven automated placement optimization.
Data Source
AI summary
Embodiments of monitoring data assets in a system to apply rules to optimize storage and access of the data assets based on data content rather than data location, by defining rules based on the monitoring attributes, wherein a rule dictates a storage location of selected data or access permissions to the data by one or more persons or groups in the system. The selected data is tagged with a defined metadata tag, and a dataset is created by running a query against a data catalog to derive the dataset. A component monitors data usage and access of data elements referenced by the dataset to detect any violations of the defined rules, and provides a notification of any violation to facilitate remedial action by a user or process. The dataset can span multiple storage devices of different types to define a single processing unit for the monitoring attributes.


