In-Process Performance Data Accumulation in Backup Data Streams

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional storage management systems face challenges in reporting performance for complex storage management jobs involving multiple data streams, as they lack the ability to reconcile performance data across components and tasks, leading to error-prone and resource-intensive reconciliation processes.

Innovation Solution

A streamlined approach where each data stream is individually tracked by data agents and media agents, generating performance reports in-process by accumulating performance data packets, which are embedded within the data stream, allowing for hierarchical analysis and eliminating the need for post-processing reconciliation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional job-level reporting is used for storage management systems, then component-specific performance logging is simplified, but performance analysis for multi-stream jobs becomes inaccurate and requires complex reconciliation processes

Engineering Contradiction:
Improveperformance analysis accuracyVSAvoidreconciliation process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments performance tracking to the data stream level rather than job level. Each data stream is assigned a unique identifier and tracked independently through the storage management system. This segmentation allows precise measurement of performance metrics for individual streams while eliminating the need to reconcile multiple component logs, as each stream carries its own performance data throughout the pipeline.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces data stream performance packets as intermediary carriers that transport performance metrics through the system. These packets are embedded with data stream identifiers and performance metrics, allowing seamless tracking across data agents, media agents, and storage devices without requiring centralized reconciliation. The packets act as self-contained mediators that preserve performance data integrity across component boundaries.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If performance data is collected from multiple components and tasks, then comprehensive performance coverage is achieved, but reconciliation of performance data becomes error-prone and resource-intensive

Engineering Contradiction:
Improveperformance data completenessVSAvoidreconciliation efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent performs preliminary action by embedding performance data packets with unique data stream identifiers at the source before data processing begins. This preliminary tagging ensures that performance metrics are associated with the correct data stream from the outset, eliminating the need for error-prone post-processing reconciliation. The performance data is collected comprehensively from multiple components but remains automatically correlated through the pre-established identifiers.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms where performance data packets circulate through the system with each component adding or verifying metrics. The system continuously monitors and updates performance data along the data stream path, providing real-time feedback on processing status. This feedback loop ensures comprehensive data collection while maintaining accuracy through continuous verification at each stage rather than single-point reconciliation.

Inventive Principle:
Principle #23Feedback

3Ease of operation

If tasks lack awareness of their hierarchical position, then task independence is maintained, but performance data reconciliation across processes and subtasks becomes complex

Engineering Contradiction:
Improvetask independenceVSAvoidhierarchical reconciliation complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent applies nesting by embedding hierarchical position information within the performance data packets themselves. Each packet contains nested identifiers that reflect the hierarchical relationship between parent processes and child subtasks. This nested structure allows tasks to remain independent and unaware of hierarchy while the performance data automatically captures the hierarchical context through embedded identifiers, eliminating the need for complex reconciliation algorithms.

Inventive Principle:
Principle #7Nested doll (Nesting)

Solution Approach 2:

The patent adds a new dimension to performance tracking by incorporating hierarchical position as an additional attribute within the performance data packets. Rather than requiring tasks to understand or communicate hierarchical relationships, the system adds this dimension of information passively through automated tracking. Each performance packet carries multi-dimensional information including stream identifier, component location, and hierarchical position, allowing comprehensive analysis without increasing operational complexity for individual tasks.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12147312B2Incrementally accumulating in-process performance data into a data stream in a secondary copy operation
Publication Date: 2024.11.19 COMMVAULT SYSTEMS INC
  • US12147312B2 patent drawing
  • US12147312B2 patent drawing
  • US12147312B2 patent drawing

AI summary

Each data stream in a backup job is individually tracked by data agent(s) and media agent(s) in its path, generating performance data packets in-process and merging them into the processed data stream. The data stream thus incrementally accumulates performance data packets from any number of successive backup processes. The in-process tracking also captures hierarchical relationships among backup processes and in-process subtending tasks, so that the resulting performance report can depict parent and child operations. The hierarchical relationships are embedded into the performance data packets and may be analyzed by parsing the data stream. The media agent transfers the data packets belonging to the secondary copy to secondary storage. The media agent analyzes the performance data packets in the data stream and generates a performance report, which covers the data stream from source to destination, based on the accumulated information carried by the performance data packets. The media agent illustratively stores the performance report locally as a flat file.