Data Labeling for Granular Backup Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current backup software systems do not adequately support data labeling on a file level, leading to inefficient and non-granular control of data sets during backup operations, as they treat all data uniformly without distinguishing between different types of data.
Innovation Solution
Implementing a data labeling process that identifies data characteristics and assigns labels, allowing for granular control by integrating a Data Label Rules Engine (DLRE) to enforce policies based on data labels, enabling more precise management of backup operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If backup software treats all data uniformly without data labeling, then the system is simple to operate, but data management precision and control granularity are insufficient
Solution Approach 1:
The patent segments data into labeled categories (e.g., confidential, public, top secret) and applies different backup policies to each category. This allows precise data management by dividing the data set into manageable segments with distinct characteristics and requirements.
Solution Approach 2:
The patent assigns different qualities or properties to different parts of the data based on labels. Each data category receives tailored backup policies, compression settings, and storage allocations appropriate to its specific requirements, rather than applying uniform treatment to all data.
2Adaptability or versatility
If backup software implements file-level data labeling, then data control granularity is improved, but the complexity of backup operations increases
Solution Approach 1:
The patent performs data labeling and classification in advance of the actual backup operation. By pre-tagging data with labels during normal operations or during initial scans, the system prepares data for differentiated backup handling without adding complexity to the execution phase of backups.
Solution Approach 2:
The system automatically discovers and applies labels to data based on predefined criteria and patterns. This self-service approach eliminates the need for manual labeling by operators, maintaining ease of operation while achieving fine-grained data control through automated classification.
3Reliability
If backup software applies uniform backup processes to all data, then the backup process is simple to implement, but data protection effectiveness varies for different data types
Solution Approach 1:
The patent applies different backup policies, retention periods, and protection measures to different data categories based on their labels. Confidential data receives enhanced protection and longer retention, while public data uses standard procedures, optimizing reliability for each data type's specific needs.
Solution Approach 2:
The patent changes key backup parameters (such as retention period, storage location, compression level) based on data labels. This allows the backup process to adapt its behavior dynamically according to data sensitivity and requirements, improving overall protection effectiveness.
Data Source
AI summary
Embodiments for a method performing data migration such as backups and restores in a network by identifying characteristics of data in a data saveset to separate the data into defined types based on respective characteristics, assigning a data label to each defined type, defining migration rules for each data label, discovering assigned labels during a migration operation; and applying respective migration rules to labeled data in the data saveset. The migration rules can dictate storage location, access rights, replication periods, retention periods, and similar parameters.


