Metadata-Based Data Backup Prioritization Across Storage Tiers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data backup systems lack an efficient method to prioritize and optimize the backup process based on the importance of data objects, leading to suboptimal storage and retrieval efficiency.
Innovation Solution
A system that determines backup scores for data objects using metadata, such as relationship and change data, to prioritize backups, storing high-priority data in faster and more reliable storage while less important data is stored in slower or less expensive media, and adjusting backup frequencies and locations accordingly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If all data objects are backed up with the same priority and frequency, then backup simplicity is maintained, but storage efficiency and retrieval speed are reduced
Solution Approach 1:
The patent segments data objects into different priority levels (e.g., critical, important, optional) based on their importance to the system. This allows the backup system to treat different data differently, backing up critical data more frequently and with higher priority while reducing resources allocated to less important data, thereby improving overall backup efficiency without treating all data uniformly
Solution Approach 2:
The patent applies local quality by assigning different backup characteristics to different data objects based on their individual importance. Each data object or group of data objects receives customized backup treatment (frequency, priority, storage location) according to its specific characteristics, rather than applying a uniform backup strategy to all data
2Reliability
If high-priority data is stored in faster and more reliable storage media, then data accessibility and reliability are improved, but storage cost increases
Solution Approach 1:
The patent implements local quality by storing different types of data on different storage media based on their priority levels. Critical data is stored on faster, more reliable storage media (e.g., SSDs, memory), while less important data is stored on slower, less expensive media (e.g., HDDs, tape). This differentiated storage approach improves reliability for critical data without requiring all data to occupy expensive storage capacity
Solution Approach 2:
The patent segments the storage infrastructure into different storage tiers (fast/expensive and slow/cheap) and assigns data objects to appropriate tiers based on their backup priority. This segmentation allows the system to optimize the balance between reliability and storage cost by placing only the necessary data on expensive media
3Reliability
If backup frequency is increased for all data objects, then data freshness and recoverability are improved, but processing time and resource consumption increase
Solution Approach 1:
The patent segments backup operations into different frequency schedules based on data importance. Critical data objects are backed up frequently (e.g., continuously or hourly), while less important data is backed up less frequently (e.g., daily or weekly). This segmented approach ensures data freshness for critical systems without unnecessarily consuming processing time and resources for all data objects
4Productivity
If automated prioritization using metadata is implemented, then backup optimization is improved, but system complexity and processing overhead increase
Solution Approach 1:
The patent implements self-service by having the system automatically determine data priority and backup requirements based on metadata and predefined rules, without requiring manual intervention. The backup system autonomously analyzes data characteristics, applies prioritization logic, and configures backup strategies automatically, reducing the need for user configuration while optimizing backup performance
Data Source
AI summary
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for backing up data. One of the methods includes, as part of a data backup process for a plurality of data objects, determining, for each of two or more data objects from the plurality of data objects and using metadata for the respective data object, a backup score that indicates an importance of the respective data object; determining different priority levels for two or more subsets of the plurality of data objects based at least in part on the backup scores; and backing up at least one of the subsets of the plurality of data objects according to the respective priorities of the subsets.

