Cloud Data Backup Snapshots with Deduplication and Controlled Restoration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data management systems lack the capability to provide fine-granularity backups and quick recovery of user data across different cloud-based applications and platforms, leading to potential data loss due to end-user errors, and fail to offer efficient search and restoration functionalities.
Innovation Solution
A data management and storage system that captures and stores snapshots of user data at hourly or five-minute intervals, allowing for fine-granularity backups, quick recovery, and aggregated search across multiple cloud platforms, utilizing a cloud-based data management application to orchestrate data management tasks, including deduplication and controlled restoration of sensitive information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If snapshots are captured at fine-granularity intervals (hourly or five-minute intervals), then data recovery precision and capability are improved, but storage requirements and system complexity increase
Solution Approach 1:
The system segments data management into distinct components: snapshot capture module, deduplication module, and restoration module. Each component operates independently at fine-granularity intervals, allowing precise data recovery without requiring the entire system to complexly coordinate all functions simultaneously.
Solution Approach 2:
The system creates multiple copies of data snapshots at different time intervals (hourly and five-minute snapshots). These copies are stored with deduplication to reduce redundancy while maintaining the ability to restore any specific point in time, thus achieving high recovery precision without proportionally increasing storage complexity.
2Reliability
If multiple snapshots are stored for fine-granularity backups, then data protection capability is improved, but storage space consumption increases
Solution Approach 1:
The system merges multiple snapshot versions by identifying and eliminating duplicate data blocks across different time points. Deduplication technology combines redundant information from hourly and five-minute snapshots into a single stored copy, maintaining complete data protection capability while significantly reducing total storage space requirements.
Solution Approach 2:
The system discards redundant duplicate data blocks from multiple snapshots while maintaining the ability to recover any lost or deleted information. By identifying and removing duplicates, the system preserves only essential unique data, thus protecting data integrity without proportionally increasing storage consumption.
3Reliability
If snapshots are captured frequently at fine-granularity intervals, then data loss prevention is improved, but processing time and resource consumption increase
Solution Approach 1:
The system implements periodic snapshot capture at two different intervals: hourly snapshots for comprehensive data protection and five-minute snapshots for critical data changes. This periodic action at varying frequencies prevents data loss while optimizing processing time by not continuously capturing snapshots at the finest granularity for all data.
Solution Approach 2:
The system applies fine-granularity five-minute snapshotting only when necessary for critical data protection, while using coarser hourly intervals for less critical data. This partial application of fine-granularity capture prevents unnecessary processing time consumption while maintaining adequate data loss prevention across all data types.
4Quantity of substance
If deduplication is implemented across multiple snapshots, then storage efficiency is improved, but computational overhead and processing complexity increase
Solution Approach 1:
The system creates simplified copies of snapshot data with metadata identifiers that track duplicate blocks across different time points. By using copy-on-write techniques and block-level deduplication, the system achieves high storage efficiency without requiring complex real-time processing of entire snapshot datasets.
Solution Approach 2:
The deduplication system serves multiple functions simultaneously: it reduces storage space, enables faster snapshot creation, and facilitates quicker restoration operations. This multi-functionality justifies the processing complexity by delivering multiple benefits from a single deduplication infrastructure.
Data Source
AI summary
Methods and systems for improving data back-up, recovery, and search across different cloud-based applications, services, and platforms are described. A data management and storage system may direct compute and storage resources within a customer's cloud-based data storage account to back-up and restore data while the customer retains full control of their data. The data management and storage system may direct the compute and storage resources within the customer's cloud-based data storage account to generate and store secondary layers that are used for generating search indexes, to generate and store shared space layers and user specific layers to facilitate the deduplication of email attachments and text blocks, to perform a controlled restoration of email snapshots such that sensitive information (e.g., restricted keywords) located within stored snapshots remains protected, and to detect and preserve emails that were received or transmitted and then deleted between two consecutive snapshots.


