Distributed Backup Metadata Aggregation for Storage Servers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional data backup processes across multiple storage servers over wide area networks, such as the Internet, are slow and expensive due to the need to transfer large data sets, which can be mitigated by generating backup data and metadata locally on each storage server without transferring the underlying data across the network.
Innovation Solution
A method where backup data is generated on each storage server independently, with backup metadata being copied and merged into a single combined metadata for quick file location and restoration, allowing backup and restoration processes to occur without transferring data across the network.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If data is transferred to a centralized server for backup, then backup metadata can be generated centrally, but data transfer time and network costs increase significantly
Solution Approach 1:
The patent segments the centralized backup process into distributed operations. Each storage server independently generates its own backup data and metadata locally, eliminating the need to transfer large datasets across the network. The segmentation allows parallel processing of backup operations on multiple servers simultaneously, dramatically reducing total backup time and network costs.
Solution Approach 2:
The patent introduces a metadata server as an intermediary that collects and merges metadata from multiple storage servers. This mediator enables centralized management of backup information without requiring transfer of actual data, resolving the contradiction between centralized control and distributed data processing.
2Reliability
If data is transferred across wide area networks for backup, then centralized backup is achieved, but network costs and transfer speed limitations worsen
Solution Approach 1:
The backup process is segmented so that each storage server performs local backup operations independently. This eliminates expensive wide area network data transfers while maintaining complete backup coverage across all servers. Only small metadata records are transferred to the metadata server, dramatically reducing network costs.
Solution Approach 2:
The patent creates local copies of backup data on each storage server and generates corresponding metadata copies. These local copies eliminate the need to transfer data across wide area networks, while the metadata copies enable centralized tracking and management of all backups through the metadata server.
3Loss of time
If backup data is generated locally on each server, then data transfer is eliminated, but coordination of synchronous backup across servers becomes complex
Solution Approach 1:
The metadata server provides feedback mechanisms that track and coordinate backup operations across multiple storage servers. Each server reports its backup status to the metadata server, which maintains a unified view of all backups and enables synchronous backup coordination without complex peer-to-peer communication between servers.
Solution Approach 2:
The metadata server acts as an intermediary that simplifies synchronization coordination. Instead of requiring complex direct communication between all pairs of storage servers, the metadata server centralizes the coordination logic, receiving backup status information from each server and managing the synchronous backup process.
4Ease of operation
If metadata is merged from multiple servers, then quick file location is enabled, but metadata collection and merging time increases
Solution Approach 1:
The metadata system is designed to be dynamic, allowing incremental updates and merges rather than requiring complete metadata collection at once. The metadata server can process and merge metadata records progressively as they become available from different storage servers, reducing the overall time required while maintaining search capability.
Data Source
AI summary
A method for generating data backups, that includes receiving, by a local backup manager executing on a local storage server, a command to initiate a backup process for a virtual data pool, and in response to receiving the command, identifying a plurality of locations pointing to a plurality of data, respectively, making a first determination that a first location points to a remote data stored on a remote storage server, in response to the first determination, sending a second command, to the remote storage server, to generate a remote backup data of the remote data, making a second determination that a second location points to a local data stored on the local storage server, and in response to the second determination, generating a local backup data of the local data, where the plurality of data backups includes the remote backup data and the local backup data.


