NDMP Multi-Stream Backup Indexing for File Server Performance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Network Data Management Protocol (NDMP) based streaming backup operations from file servers or clusters are inefficient due to lack of native multi-streaming capabilities, leading to prolonged backup times and potential failures, especially with large data sets.
Innovation Solution
Implementing NDMP granular multi-streaming by activating multiple concurrent data streams based on specific parameters, taking snapshots of source volumes, and using a proprietary index for optimized data stream allocation and granular backup and restore operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If NDMP protocol is used for backup operations from file servers, then backup reliability is improved, but backup speed deteriorates due to lack of native multi-streaming capabilities
Solution Approach 1:
The patent segments the backup data into multiple distinct data streams, each handled by a separate NDMP session. Instead of using a single backup stream, the system creates multiple parallel streams that can be backed up simultaneously, thereby increasing overall backup speed while maintaining the reliability of the NDMP protocol for each individual stream.
Solution Approach 2:
The patent introduces a new dimension of parallelism by running multiple NDMP sessions concurrently. Rather than improving speed within a single sequential NDMP stream, the system adds temporal and spatial parallelism by establishing multiple independent backup channels, effectively transforming the backup process from a single-dimensional sequential operation to a multi-dimensional parallel operation.
2Adaptability or versatility
If single data stream backup is used with NDMP, then protocol compatibility is maintained, but productivity deteriorates due to prolonged backup times
Solution Approach 1:
The backup data is segmented into multiple independent streams, each compatible with the NDMP protocol. This segmentation allows the system to maintain full NDMP protocol compatibility for each stream while dramatically improving productivity through parallel processing of multiple streams simultaneously.
Solution Approach 2:
The patent enables continuous backup operations by running multiple NDMP sessions in parallel without interruption. While a single NDMP stream must complete sequentially, multiple concurrent streams maintain continuous useful action across all data sources, significantly boosting overall backup productivity while preserving protocol compatibility.
3Reliability
If backup operations are performed on large data sets with single stream, then data completeness is ensured, but loss of time increases due to re-running failed jobs
Solution Approach 1:
Large data sets are segmented into multiple manageable streams, each of which can be backed up independently. If one stream fails, only that specific stream needs to be re-run rather than the entire backup job, thereby reducing time loss while ensuring data completeness through parallel execution of multiple streams.
Solution Approach 2:
The patent applies partial action by allowing individual streams to be completed or failed independently. Instead of requiring complete success of all streams to consider the backup successful, the system allows partial completion and enables selective re-execution of only failed streams, reducing overall backup time and minimizing loss of time while maintaining data completeness.
4Manufacturing precision
If granular backup copies are created for each directory, then restore granularity is improved, but device complexity increases due to multiple backup copies
Solution Approach 1:
The backup system creates segmented backup copies at the directory level, allowing granular restore operations. Each directory or file can be restored independently from its corresponding backup copy, providing fine-grained restore precision. The system manages this complexity through automated bookkeeping that tracks which backup copy corresponds to which data source.
Solution Approach 2:
The patent introduces an intermediary indexing mechanism that maps data sources to their corresponding backup copies. This intermediary layer manages the complexity of multiple granular backup copies by providing a unified interface for backup and restore operations, thereby enabling fine-grained restore capability without exposing the full complexity of managing multiple individual backup copies to the user.
Data Source
AI summary
Multiple substantially concurrent data streams with NDMP protocol improve robustness, performance, and granularity of backup and restore operations from/to a filer. NDMP data streams are initially allocated based on inventorying the root level of each filer volume. A best effort to balance the multiple NDMP data streams allocates them based on data amounts used in each volume. Orphaned files are also collected and backed up. Subsequent full backup jobs leverage a proprietary index generated in preceding full backup jobs to obtain better performance and to better balance the NDMP data streams by creating substantially co-equal groupings of source data. The index comprises granular information which is not available from querying the filer. The size of each individual backup copy from a preceding full backup job and/or the size of subtending subdirectories or individual backed up files therein is used by later backup jobs to fine tune NDMP data stream allocation.


