Media Agent Segregation for Deduplication Interoperability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional data storage management systems face challenges in efficiently managing and protecting data due to high processing loads and lack of interoperability between components and deduplication appliances, leading to increased resource utilization and complexity in backup, restore, and tertiary copy operations.
Innovation Solution
A tiered indexing approach is implemented to minimize data retention and processing loads, where media agents segregate data streams and utilize deduplication appliances for efficient data handling, deduplication, and storage, while retaining indexing and tracking capabilities within the data storage management system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional data storage management systems process and prepare all data copies locally, then data protection reliability is improved, but processing load and resource utilization increase significantly
Solution Approach 1:
The patent segments the data management system into distinct components: media agents that handle indexing and tracking, and deduplication appliances that handle processing and storage. This segmentation allows each component to specialize in specific tasks, reducing the processing load on media agents while maintaining data protection reliability through coordinated operation.
Solution Approach 2:
The patent introduces deduplication appliances as intermediary components between media agents and storage systems. These appliances act as mediators that receive data from media agents, perform deduplication processing, and store the processed data, thereby reducing the processing burden on media agents while ensuring data protection.
2Productivity
If data is retained and stored at media agents for indexing and tracking, then indexing efficiency is improved, but data storage requirements and system complexity increase
Solution Approach 1:
The patent extracts the heavy data storage function from media agents and places it at deduplication appliances. Media agents retain only essential indexing and tracking information, while the actual data copies are stored at deduplication appliances. This extraction reduces data retention requirements at media agents while maintaining indexing efficiency through centralized data storage.
3Ease of operation
If media agents handle all data processing for tertiary copies and restore operations, then operational control is improved, but processing time and resource utilization increase
Solution Approach 1:
The patent segments tertiary copy and restore operations between media agents and deduplication appliances. Media agents handle high-level coordination and control, while deduplication appliances perform the actual data processing and retrieval. This segmentation reduces processing time by distributing computational tasks while maintaining operational control through media agent coordination.
4Quantity of substance
If deduplication is performed on all data streams, then storage efficiency is improved, but processing load and complexity increase
Solution Approach 1:
The patent applies deduplication selectively rather than uniformly to all data streams. Deduplication appliances perform deduplication on payload data streams where it provides maximum benefit, while other data streams are handled differently. This local quality approach improves storage efficiency for the most suitable data types while reducing overall processing complexity.
Data Source
AI summary
Illustrative storage manager and media agent are enhanced to interoperate with deduplication appliances. Advantages are realized when making secondary and tertiary copies and also when restoring from a deduplication appliance. Tiered indexing minimizes how much data is retained and stored at media agents. Tiered indexing enables media agents to efficiently extract needed information from deduplication appliances to make tertiary copies and to restore backed up copies. Interoperability techniques include media agents generating separate data streams to the deduplication appliance. Each data stream carries a different kind of data, e.g., payload data, metadata content, or high-level index information. On initial backup, the media agent instructs the deduplication appliance to deduplicate the payload data stream but not the other data streams, thus intelligently applying resources to data most likely to benefit from deduplication. For tertiary copies (copies of pre-existing copies at the deduplication appliance), the media agent avoids handling payload data altogether.


