Differential Backup System Using Data Chunking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data backup and archiving methods are inefficient and costly, as they redundantly store identical data, lack scalability, and require significant administrative overhead, making it impractical for large organizations to manage increasing data storage needs.

Innovation Solution

A distributed, differential electronic-data backup and archiving system that employs chunking methods to identify and store only unique data chunks, allowing for efficient storage and retrieval by collocating data objects with shared chunks within a single cell, reducing physical storage requirements and administrative burdens.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional backup methods store complete copies of all data, then data reliability is ensured, but storage space requirements increase significantly

Engineering Contradiction:
Improvedata reliabilityVSAvoidstorage space
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent segments data into unique chunks and stores only distinct segments across the network. Each file is broken down into chunked objects that can be independently stored and referenced, eliminating redundant storage of identical data portions while maintaining complete data reliability through distributed chunk storage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Instead of copying entire files, the system creates references to unique data chunks. Multiple files that share common content reference the same chunked objects, reducing storage requirements while ensuring data can be fully reconstructed when needed.

Inventive Principle:
Principle #26Copying

2Quantity of substance

If more servers and mass-storage devices are purchased to manage increasing data, then storage capacity increases, but expense and administrative overhead increase

Engineering Contradiction:
Improvestorage capacityVSAvoidadministrative overhead
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent creates a universal chunked object storage system where the same infrastructure serves multiple purposes: backup, archiving, and data retrieval. The distributed network of personal computers with hard drives functions as both storage media and processing nodes, eliminating the need for specialized backup hardware and reducing administrative complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system enables automatic data chunking, storage, and retrieval operations without requiring manual intervention. The distributed network automatically manages data segmentation, chunk storage, and reconstruction, reducing the need for specialized backup administrators and simplifying data management tasks.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If data is distributed across multiple cells, then scalability improves, but data retrieval complexity increases

Engineering Contradiction:
ImprovescalabilityVSAvoiddata retrieval complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a reference counting mechanism that provides feedback about chunk usage across the distributed network. This feedback system tracks which chunks are referenced by which files, enabling efficient data retrieval by identifying all locations of required chunks and coordinating their assembly, thus managing retrieval complexity in scalable systems.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system pre-computes and stores metadata about data chunk locations and references in a distributed manner. This preliminary organization of chunk information across cells enables efficient retrieval operations by providing a map of where data segments are stored, reducing the complexity of locating and assembling distributed data objects.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8862841B2Method and system for scaleable, distributed, differential electronic-data backup and archiving
Publication Date: 2014.10.14 HEWLETT PACKARD ENTERPRISE DEV LP
  • US8862841B2 patent drawing
  • US8862841B2 patent drawing
  • US8862841B2 patent drawing

AI summary

One embodiment of the present invention provides a distributed, differential electronic-data backup and archiving system that includes client computers and cells. Client computers execute front-end-application components of the distributed, differential electronic-data backup and archiving system, the front-end application components receiving data objects from client computers and sending the received data objects to cells of the distributed, differential electronic-data backup and archiving system for storage. Cells within the distributed, differential electronic-data backup and archiving system store the data objects, each cell comprising at least one computer system with attached mass-storage and each cell storing entire data objects as lists that reference stored, unique data chunks within the cell, a cell storing all of the unique data chunks for all data objects stored in the cell.