Backup Policy Alignment via Data Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data classification methods do not effectively analyze and prioritize data in the data protection workflow, leading to increased backup times and costs due to the proprietary nature of storage formats and lack of consideration for data sensitivity and incremental behavior, resulting in data debris accumulation.
Innovation Solution
Harnessing critical parameters from infrastructure configurations and data relevance metrics to generate an improved data protection workflow that aligns backup policies with business criticality, enabling automated data removal, one-time archival, and cost reduction through controller-based replication.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional backup policy scans all files to check for updates, then data protection coverage is comprehensive, but backup time is greatly increased
Solution Approach 1:
The patent extracts and analyzes file metadata (keywords, sensitivity levels, modification dates) separately from full file content scanning. By taking out only the essential classification information and using it to identify candidate files for backup, the system maintains comprehensive protection coverage while dramatically reducing the time spent scanning all files.
Solution Approach 2:
The system performs preliminary data classification and sensitivity analysis on files before the actual backup process. By pre-identifying sensitive files and their priority levels through metadata analysis, the backup policy can focus subsequent scanning and backup operations only on relevant files, reducing overall backup time while maintaining comprehensive protection.
2Reliability
If conventional backup policy backs up all updated files, then data protection is thorough, but data center storage costs increase
Solution Approach 1:
The patent applies different backup retention policies to different files based on their sensitivity classification. Sensitive files receive full backup protection with longer retention periods, while non-sensitive files use incremental backups with shorter retention. This local differentiation of backup quality based on file importance reduces overall storage consumption while maintaining thorough protection for critical data.
Solution Approach 2:
The system automatically identifies and removes obsolete backup data based on file sensitivity levels and backup policies. By discarding redundant backups of non-sensitive files and retaining only essential backups of sensitive files, the system maintains thorough data protection for critical information while reducing total storage volume through selective data elimination.
3Device complexity
If data classification analyzes only operational data, then processing simplicity is maintained, but backup system clutter increases due to unanalyzed backup data
Solution Approach 1:
The patent extends the data classification system to handle both operational data and backup data using the same classification framework. The classification engine analyzes metadata from both data types, applying consistent sensitivity rules and keywords to identify important files regardless of whether they are currently in use or stored as backups, thereby preventing backup system clutter while maintaining processing consistency.
4Adaptability or versatility
If file-by-file network based backup is used, then backup flexibility is high, but backup stream costs increase
Solution Approach 1:
The patent merges multiple individual file backup operations into consolidated backup streams based on sensitivity classification. Files with the same sensitivity level and backup requirements are grouped together and backed up in batch operations rather than individual file-by-file transfers. This reduces the total number of network streams and associated costs while maintaining the flexibility to handle different file types and priorities.
Data Source
AI summary
A backup and archival policy method, system, and non-transitory computer readable medium, includes harnessing of metrics of data classification including both operational data and backup data from an end-to-end stack from a backup Information Lifecycle Governance (ILM) viewpoint.


