Differential Backup System Using Data Chunking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data backup and archiving methods are inefficient and costly, as they redundantly store identical data, lack scalability, and require significant administrative overhead, making it impractical for large organizations to manage increasing data storage needs.
Innovation Solution
A distributed, differential electronic-data backup and archiving system that employs chunking methods to identify and store only unique data chunks, allowing for efficient storage and retrieval by collocating data objects with shared chunks within a single cell, reducing physical storage requirements and administrative burdens.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional backup methods store complete copies of all data, then data reliability is ensured, but storage space requirements increase significantly
Solution Approach 1:
The patent segments data into unique chunks and stores only distinct segments across the network. Each file is broken down into chunked objects that can be independently stored and referenced, eliminating redundant storage of identical data portions while maintaining complete data reliability through distributed chunk storage.
Solution Approach 2:
Instead of copying entire files, the system creates references to unique data chunks. Multiple files that share common content reference the same chunked objects, reducing storage requirements while ensuring data can be fully reconstructed when needed.
2Quantity of substance
If more servers and mass-storage devices are purchased to manage increasing data, then storage capacity increases, but expense and administrative overhead increase
Solution Approach 1:
The patent creates a universal chunked object storage system where the same infrastructure serves multiple purposes: backup, archiving, and data retrieval. The distributed network of personal computers with hard drives functions as both storage media and processing nodes, eliminating the need for specialized backup hardware and reducing administrative complexity.
Solution Approach 2:
The system enables automatic data chunking, storage, and retrieval operations without requiring manual intervention. The distributed network automatically manages data segmentation, chunk storage, and reconstruction, reducing the need for specialized backup administrators and simplifying data management tasks.
3Adaptability or versatility
If data is distributed across multiple cells, then scalability improves, but data retrieval complexity increases
Solution Approach 1:
The patent implements a reference counting mechanism that provides feedback about chunk usage across the distributed network. This feedback system tracks which chunks are referenced by which files, enabling efficient data retrieval by identifying all locations of required chunks and coordinating their assembly, thus managing retrieval complexity in scalable systems.
Solution Approach 2:
The system pre-computes and stores metadata about data chunk locations and references in a distributed manner. This preliminary organization of chunk information across cells enables efficient retrieval operations by providing a map of where data segments are stored, reducing the complexity of locating and assembling distributed data objects.
Data Source
AI summary
One embodiment of the present invention provides a distributed, differential electronic-data backup and archiving system that includes client computers and cells. Client computers execute front-end-application components of the distributed, differential electronic-data backup and archiving system, the front-end application components receiving data objects from client computers and sending the received data objects to cells of the distributed, differential electronic-data backup and archiving system for storage. Cells within the distributed, differential electronic-data backup and archiving system store the data objects, each cell comprising at least one computer system with attached mass-storage and each cell storing entire data objects as lists that reference stored, unique data chunks within the cell, a cell storing all of the unique data chunks for all data objects stored in the cell.


