Content-Based Dataset Mobility for Data Center Resource Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data management systems struggle to efficiently manage the movement and lifecycle of large-scale datasets across diverse storage environments, particularly in data centers, due to the complexity of handling data spread across multiple sources and locations, leading to challenges in compliance, scalability, and resource optimization.
Innovation Solution
A dataset management system that utilizes metadata-based datasets to create logical collections of files and objects, allowing for content-based data protection and lifecycle management, enabling efficient data backup, restore, move, tier, and deletion operations, regardless of storage location, and includes a data center mobility process that analyzes dataset growth patterns to make informed relocation decisions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is moved across data centers using traditional methods, then data center capacity constraints are addressed, but the complexity of tracking and managing data locations and lifecycles increases significantly
Solution Approach 1:
The patent creates a virtual copy of the data catalog that travels with the data during migration. This virtual catalog contains all necessary metadata and lifecycle information, eliminating the need for complex tracking across multiple physical locations. The copy principle allows the system to manage data mobility without proportionally increasing management complexity.
Solution Approach 2:
The patent introduces a data catalog as an intermediary layer between the data and the management system. This catalog serves as a mediator that tracks data locations, lifecycle policies, and migration status centrally, simplifying the management of cross-data center operations by providing a single point of control.
2Reliability
If traditional location-based data protection methods are used, then data can be tracked and protected, but the system cannot efficiently handle data spread across multiple locations and sources
Solution Approach 1:
The patent implements a universal data catalog system that can track and protect data regardless of its location or source. The catalog uses content-based identification rather than location-based tracking, allowing it to universally manage data across multiple data centers, cloud environments, and storage types with a single system.
Solution Approach 2:
The patent changes the fundamental parameter for data identification from location-based (where the data is) to content-based (what the data is). This parameter change allows the system to maintain reliable protection while adapting to data distributed across multiple locations, as the catalog follows the data based on its identity rather than its physical position.
3Reliability
If manual data lifecycle management is used, then compliance rules can be applied, but the system cannot scale to handle increasing data volumes
Solution Approach 1:
The patent implements automated lifecycle management where the data catalog itself manages tracking, policy application, and compliance enforcement without manual intervention. The system serves itself by automatically monitoring data locations, applying retention policies, and managing migrations based on predefined rules, enabling scalability while maintaining compliance.
Solution Approach 2:
The patent applies lifecycle policies and compliance rules in advance through the data catalog before data issues arise. By pre-configuring retention policies, migration rules, and protection parameters in the catalog, the system automatically enforces compliance without reactive manual management, scaling efficiently with data growth.
4Productivity
If data is packaged and moved without intelligent selection, then migration can proceed, but significant effort and resources are wasted moving inappropriate data
Solution Approach 1:
The patent uses the data catalog to provide feedback on data characteristics, access patterns, and lifecycle status before migration decisions are made. This feedback mechanism allows the system to intelligently select which data should be migrated based on actual usage and compliance requirements, avoiding the waste of moving inappropriate data while maintaining efficient migration speeds.
Data Source
AI summary
Optimizing data movement from a source data center to a target data center by grouping metadata of data objects into a dataset that encompasses data processed identically by a data processing operation, where the dataset defines a single data access unit for these data objects, and the data processing operation processes them as a single unit based on data content rather than physical or logical data location of the data objects. A mobility process correlates an increase in a number of metadata elements in the dataset with a growth rate of the dataset, and compares the dataset growth rate with historical or similar dataset growth rate data. The comparison is used to determine when and where to move data from the source to the target data center based on a forecast of accelerated growth indicated by the comparing step, and a consideration of target data center resources, data move costs, and streaming data effects.


