Backup System Heuristic Data Prioritization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing backup software applications do not consider the importance or priority of data during backup operations, leading to inefficient use of resources and potential data loss, especially in urgent scenarios like natural disasters.
Innovation Solution
Implementing heuristic techniques such as social network data analysis, topic modeling, and cluster analysis to identify and prioritize data subsets for backup, allowing for expedited backup of high-priority data to faster storage devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all data is backed up using traditional backup operations, then complete data protection is achieved, but resource consumption and time required increase significantly
Solution Approach 1:
The patent segments the dataset into multiple subsets based on priority levels (first subset with high priority, second subset with low priority). This segmentation allows differential backup strategies to be applied to different portions of data, protecting critical data while reducing overall backup resource consumption.
Solution Approach 2:
The patent applies different backup qualities to different data subsets. High-priority data in the first subset receives expedited backup with higher reliability guarantees, while low-priority data in the second subset receives standard backup treatment. This local quality differentiation resolves the contradiction by ensuring complete protection where needed while optimizing efficiency where less critical.
2Reliability
If incremental backup operations are performed on all changed data, then data recovery capability is maintained, but storage resources are wasted on non-critical data
Solution Approach 1:
The patent extracts and identifies only the critical portion of changed data (first subset) that requires backup, separating it from non-critical changed data (second subset). By taking out only the essential data for backup operations, the system maintains recovery capability for important information while avoiding storage waste on insignificant data.
Solution Approach 2:
The patent applies partial action by performing complete backup operations only on the first subset of high-priority data, while applying reduced or no backup operations on the second subset of low-priority data. This partial approach maintains adequate recovery capability for critical data while reducing storage resource consumption.
3Ease of operation
If backup priority is not considered, then backup operation simplicity is maintained, but urgent data protection needs are not met
Solution Approach 1:
The patent performs preliminary analysis of the dataset to identify and prioritize critical data subsets before executing backup operations. By预先 determining which data requires expedited protection, the system can automatically apply appropriate backup strategies without requiring complex user input or manual prioritization, thus maintaining operational simplicity while addressing urgent protection needs.
Solution Approach 2:
The backup system performs self-service by automatically analyzing data priorities and applying differential backup strategies without external intervention. The system autonomously identifies high-priority subsets and expedates their backup, eliminating the need for users to manually prioritize data while still achieving time-critical protection goals.
Data Source
AI summary
Various systems, methods, and processes to analyze datasets using heuristic-based data analysis and prioritization techniques to identify, derive, and/or select subsets with important and/or high-priority data for preferential backup are disclosed. A request to perform a backup operation that identifies a dataset to be backed up to a storage device is received. A subset of data is identified and selected from the dataset by analyzing the dataset using one or more prioritization techniques. A backup operation is performed by storing the subset of data in the storage device.


