In-Process Performance Data Accumulation in Backup Data Streams
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional storage management systems face challenges in reporting performance for complex storage management jobs involving multiple data streams, as they lack the ability to reconcile performance data across components and tasks, leading to error-prone and resource-intensive reconciliation processes.
Innovation Solution
A streamlined approach where each data stream is individually tracked by data agents and media agents, generating performance reports in-process by accumulating performance data packets, which are embedded within the data stream, allowing for hierarchical analysis and eliminating the need for post-processing reconciliation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional job-level reporting is used for storage management systems, then component-specific performance logging is simplified, but performance analysis for multi-stream jobs becomes inaccurate and requires complex reconciliation processes
Solution Approach 1:
The patent segments performance tracking to the data stream level rather than job level. Each data stream is assigned a unique identifier and tracked independently through the storage management system. This segmentation allows precise measurement of performance metrics for individual streams while eliminating the need to reconcile multiple component logs, as each stream carries its own performance data throughout the pipeline.
Solution Approach 2:
The patent introduces data stream performance packets as intermediary carriers that transport performance metrics through the system. These packets are embedded with data stream identifiers and performance metrics, allowing seamless tracking across data agents, media agents, and storage devices without requiring centralized reconciliation. The packets act as self-contained mediators that preserve performance data integrity across component boundaries.
2Reliability
If performance data is collected from multiple components and tasks, then comprehensive performance coverage is achieved, but reconciliation of performance data becomes error-prone and resource-intensive
Solution Approach 1:
The patent performs preliminary action by embedding performance data packets with unique data stream identifiers at the source before data processing begins. This preliminary tagging ensures that performance metrics are associated with the correct data stream from the outset, eliminating the need for error-prone post-processing reconciliation. The performance data is collected comprehensively from multiple components but remains automatically correlated through the pre-established identifiers.
Solution Approach 2:
The patent implements feedback mechanisms where performance data packets circulate through the system with each component adding or verifying metrics. The system continuously monitors and updates performance data along the data stream path, providing real-time feedback on processing status. This feedback loop ensures comprehensive data collection while maintaining accuracy through continuous verification at each stage rather than single-point reconciliation.
3Ease of operation
If tasks lack awareness of their hierarchical position, then task independence is maintained, but performance data reconciliation across processes and subtasks becomes complex
Solution Approach 1:
The patent applies nesting by embedding hierarchical position information within the performance data packets themselves. Each packet contains nested identifiers that reflect the hierarchical relationship between parent processes and child subtasks. This nested structure allows tasks to remain independent and unaware of hierarchy while the performance data automatically captures the hierarchical context through embedded identifiers, eliminating the need for complex reconciliation algorithms.
Solution Approach 2:
The patent adds a new dimension to performance tracking by incorporating hierarchical position as an additional attribute within the performance data packets. Rather than requiring tasks to understand or communicate hierarchical relationships, the system adds this dimension of information passively through automated tracking. Each performance packet carries multi-dimensional information including stream identifier, component location, and hierarchical position, allowing comprehensive analysis without increasing operational complexity for individual tasks.
Data Source
AI summary
Each data stream in a backup job is individually tracked by data agent(s) and media agent(s) in its path, generating performance data packets in-process and merging them into the processed data stream. The data stream thus incrementally accumulates performance data packets from any number of successive backup processes. The in-process tracking also captures hierarchical relationships among backup processes and in-process subtending tasks, so that the resulting performance report can depict parent and child operations. The hierarchical relationships are embedded into the performance data packets and may be analyzed by parsing the data stream. The media agent transfers the data packets belonging to the secondary copy to secondary storage. The media agent analyzes the performance data packets in the data stream and generates a performance report, which covers the data stream from source to destination, based on the accumulated information carried by the performance data packets. The media agent illustratively stores the performance report locally as a flat file.


