Media Agent State Management via Performance Trending
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data storage management systems face challenges in detecting anomalies in event occurrences and job statuses, which can lead to overutilization of computing resources and negatively impact system performance.
Innovation Solution
The implementation of anomaly detection techniques, including time-series decomposition and alert generation, to identify deviations in event frequencies and job durations, allowing for proactive management and resource allocation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If anomaly detection techniques are implemented to detect deviations in event frequencies and job durations, then system reliability is improved, but device complexity increases
Solution Approach 1:
The system segments the monitoring process into distinct functional modules: event frequency analysis module that tracks event occurrences over time, job duration analysis module that measures completion times, and alert generation module that triggers notifications. This segmentation allows each module to specialize in specific detection tasks, improving overall reliability while managing complexity through modular design.
Solution Approach 2:
The patent introduces intermediary components including a data collection layer that gathers raw events and job information, and an analysis layer that processes this data to detect anomalies. These intermediaries buffer the complexity between data sources and decision-making processes, enabling reliable anomaly detection without directly coupling all system components.
2Productivity
If proactive management and resource allocation are enabled through anomaly detection, then productivity is improved, but use of energy increases
Solution Approach 1:
The system performs preliminary analysis of event frequencies and job durations to detect anomalies before they cause significant performance degradation. By identifying trends and deviations early, the system can proactively allocate resources or adjust operations, improving productivity while avoiding the energy costs associated with reactive problem-solving and system failures.
Solution Approach 2:
The patent dynamically adjusts monitoring parameters such as alert thresholds and analysis intervals based on system conditions. During normal operation, less intensive monitoring reduces energy consumption, while during periods of detected anomalies or high workload, the system intensifies analysis to maintain productivity. This adaptive parameter adjustment balances productivity gains with energy efficiency.
Data Source
AI summary
Described herein are techniques for automating media agent state management. For example, if a media agent is running poorly, then the media agent can be disabled and an alternate media agent can perform secondary copy job operations in place of the poorly running media agent. To determine whether a media agent is running poorly, a storage manager can determine whether the media agent has an anomalous number of failed jobs, pending jobs, and/or long running jobs and/or can determine whether the amount of resources used by the media agent is high or is increasing constantly, at a constant rate, or at a near constant rate.


