Office 365 Data Protection via On-Premises Deduplication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data protection systems for cloud-based applications like Microsoft Office365 lack efficient Point-in-Time (PIT) backup capabilities and are computationally expensive, with existing solutions relying on deduplication technology that is costly and resource-intensive.

Innovation Solution

A cost-efficient data protection system that utilizes object storage and minimal compute resources, employing a lightweight database to manage metadata and implement data tiering, allowing for efficient backup and restore operations while balancing storage and compute costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If deduplication technology is used to reduce storage costs, then storage efficiency is improved, but computational cost increases significantly

Engineering Contradiction:
Improvestorage sizeVSAvoidcomputational cost
Core Design Contradiction:
Quantity of substanceVSUse of energy by moving object

Solution Approach 1:

The patent extracts the deduplication computation from the cloud environment and relocates it to on-premises infrastructure. This allows the cloud to store compressed backup images without bearing the computational cost of creating them, effectively separating the storage function from the computation function across different locations.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an on-premises infrastructure as an intermediary between the source system and cloud storage. This intermediary performs the computationally intensive deduplication operations locally and transfers only the compressed results to the cloud, acting as a buffer that protects the cloud from high computational loads.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If cloud compute resources are used for data protection, then data protection capability is improved, but cost efficiency deteriorates

Engineering Contradiction:
Improvedata protection capabilityVSAvoidcost efficiency
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent segments the data protection workflow into distinct phases: computation-intensive deduplication operations are performed on-premises, while storage and retrieval operations are handled by cloud infrastructure. This segmentation allows each component to operate in its most cost-effective environment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates compressed copies of backup data on-premises before transferring to the cloud. By pre-compressing data locally using available compute resources, the system avoids the need for expensive cloud-based compression operations while still achieving space-efficient storage.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If Point-in-Time backup is implemented, then data recovery flexibility is improved, but system complexity increases

Engineering Contradiction:
Improvedata recovery flexibilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary deduplication and compression of backup data before it reaches the cloud. By pre-processing data on-premises and creating organized backup images with metadata, the system simplifies subsequent Point-in-Time recovery operations in the cloud, as the data is already prepared and indexed for efficient retrieval.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11669402B2Highly efficient native application data protection for office 365
Publication Date: 2023.06.06 EMC IP HLDG CO LLC
  • US11669402B2 patent drawing
  • US11669402B2 patent drawing
  • US11669402B2 patent drawing

AI summary

Embodiments for a method of storing documents using a document data protection process. Documents are first compressed and stored in a container along with selected metadata. An Document Record is created for each document. A Container Record is created for each newly created container, and a Backup Record is created for each container for each backup. Once the required records are created, the process facilitates the execution of backup operations, such as full or incremental backups of the stored documents. Data tiering is supported so that low cost object storage in the public cloud is used instead of expensive processing methods like deduplication. A user interface receives a user setting dictating a storage media storing the container based on a relative availability of the storage media versus cost of storage.