Dynamic Service Condition Escalation in Data Center Monitoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data center monitoring systems face challenges in efficiently escalating service conditions, particularly during large-scale failures or natural disasters, which can hinder communication between monitoring systems and service elements, leading to inadequate response to critical issues.
Innovation Solution
A monitoring system that dynamically escalates service conditions based on access conditions by evaluating its own access to the data center and communicating with geographically distinct monitoring systems to determine the scope of failures, initiating escalated responses for large-scale issues and non-escalated responses for localized problems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If a monitoring system automatically repairs or recovers failed service elements, then the response time to failures is reduced, but the system cannot handle large-scale failures where communication between monitoring systems and service elements is inhibited
Solution Approach 1:
The system dynamically adjusts the response strategy based on the type of failure detected. For localized failures, automated repair is initiated immediately. For large-scale failures where communication is inhibited, the system escalates to manual intervention by staff personnel, optimizing the response approach based on real-time conditions
Solution Approach 2:
The monitoring system acts as an intermediary that detects failure patterns and determines whether to initiate automated repair or escalate to human staff. It mediates between the automated repair capability and manual intervention based on the scope and nature of the detected failure
2Reliability
If staff personnel are notified for all service element failures, then comprehensive coverage of failures is achieved, but unnecessary notifications are generated for failures that can be handled automatically
Solution Approach 1:
Different notification strategies are applied to different types of failures. Localized failures that can be automatically repaired do not trigger staff notifications, while large-scale failures that require manual intervention do trigger notifications. This localizes the notification quality to match the specific failure condition
3Productivity
If automated repair operations are initiated for all failures, then response efficiency is improved, but failures requiring manual attention may be mishandled
Solution Approach 1:
The system dynamically selects the appropriate repair approach based on failure analysis. Automated repair is initiated for localized failures where it is appropriate and effective. For large-scale failures where communication is inhibited or manual intervention is required, the system escalates to staff personnel, ensuring the right response method is applied to each failure type
Data Source
AI summary
Systems, methods, and software are provided for dynamically escalating service conditions associated with data center failures. In one implementation, a monitoring system detects a service condition. The service condition may be indicative of a failure of at least one service element within a data center monitored by the monitoring system. The monitoring system determines whether or not the service condition qualifies for escalation based at least in part on an access condition associated with the data center. The access condition may be identified by at least another monitoring system that is located in a geographic region distinct from that of the first monitoring system. Upon determining that the service condition qualifies for escalation, the monitoring system escalates the service condition to an escalated condition and initiates an escalated response.


