Office 365 Data Protection via On-Premises Deduplication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data protection systems for cloud-based applications like Microsoft Office365 lack efficient Point-in-Time (PIT) backup capabilities and are computationally expensive, with existing solutions relying on deduplication technology that is costly and resource-intensive.
Innovation Solution
A cost-efficient data protection system that utilizes object storage and minimal compute resources, employing a lightweight database to manage metadata and implement data tiering, allowing for efficient backup and restore operations while balancing storage and compute costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If deduplication technology is used to reduce storage costs, then storage efficiency is improved, but computational cost increases significantly
Solution Approach 1:
The patent extracts the deduplication computation from the cloud environment and relocates it to on-premises infrastructure. This allows the cloud to store compressed backup images without bearing the computational cost of creating them, effectively separating the storage function from the computation function across different locations.
Solution Approach 2:
The patent introduces an on-premises infrastructure as an intermediary between the source system and cloud storage. This intermediary performs the computationally intensive deduplication operations locally and transfers only the compressed results to the cloud, acting as a buffer that protects the cloud from high computational loads.
2Reliability
If cloud compute resources are used for data protection, then data protection capability is improved, but cost efficiency deteriorates
Solution Approach 1:
The patent segments the data protection workflow into distinct phases: computation-intensive deduplication operations are performed on-premises, while storage and retrieval operations are handled by cloud infrastructure. This segmentation allows each component to operate in its most cost-effective environment.
Solution Approach 2:
The patent creates compressed copies of backup data on-premises before transferring to the cloud. By pre-compressing data locally using available compute resources, the system avoids the need for expensive cloud-based compression operations while still achieving space-efficient storage.
3Adaptability or versatility
If Point-in-Time backup is implemented, then data recovery flexibility is improved, but system complexity increases
Solution Approach 1:
The patent performs preliminary deduplication and compression of backup data before it reaches the cloud. By pre-processing data on-premises and creating organized backup images with metadata, the system simplifies subsequent Point-in-Time recovery operations in the cloud, as the data is already prepared and indexed for efficient retrieval.
Data Source
AI summary
Embodiments for a method of storing documents using a document data protection process. Documents are first compressed and stored in a container along with selected metadata. An Document Record is created for each document. A Container Record is created for each newly created container, and a Backup Record is created for each container for each backup. Once the required records are created, the process facilitates the execution of backup operations, such as full or incremental backups of the stored documents. Data tiering is supported so that low cost object storage in the public cloud is used instead of expensive processing methods like deduplication. A user interface receives a user setting dictating a storage media storing the container based on a relative availability of the storage media versus cost of storage.


