Automated Server Outage Data Management via Agent Retrieval
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current ITIL-compliant systems rely on manual data entry for tracking server outages, leading to inaccurate records and difficulties in determining the cost and impact of server outages, especially when metrics are based on user-generated data.
Innovation Solution
An automated system that retrieves outage data from servers, using agents to access event logs and determine the number of users affected, allowing for the creation of accurate outage records and cost calculations, which are then linked to incident records for comprehensive monitoring.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual data entry is used to track server outages according to ITIL compliance, then the system maintains procedural correctness and creates incident/problem/error records, but the data accuracy deteriorates and becomes unreflective of actual outage conditions
Solution Approach 1:
The monitored server automatically generates outage event data and transmits it to the monitoring application without requiring manual data entry. The server itself serves as the data source, eliminating the need for users to manually create incident records and ensuring the data accurately reflects actual outage conditions.
Solution Approach 2:
The manual mechanical process of users creating incident/problem/error records is replaced by an automated electronic system. The monitoring application automatically receives outage event data from the server, determines outage duration, and generates structured outage records, replacing the manual ITIL compliance process with an automated electronic data collection system.
2Ease of operation
If manual user-generated data is used for outage tracking, then the system can create incident records, but the ability to accurately determine outage cost and impact deteriorates
Solution Approach 1:
The monitoring application receives automated feedback from the monitored server about outage events and automatically processes this data to determine outage duration and impact. The system uses server data such as the number of users affected to calculate outage cost, providing accurate feedback loops that eliminate information loss.
Solution Approach 2:
The monitoring application acts as an intermediary between the monitored server and the outage record-keeping system. It receives raw outage event data from the server, processes it to determine outage duration and impact, and generates structured outage records that accurately reflect the actual cost and impact of outages.
3Measurement precision
If automated outage data collection is implemented, then data accuracy and outage cost determination improve, but the system complexity increases
Solution Approach 1:
The monitoring application performs multiple functions: it receives outage event data from servers, determines outage duration by comparing event timestamps, calculates outage impact using server data, and generates structured outage records. This multi-functionality consolidates what would otherwise require separate systems into a single automated platform.
Solution Approach 2:
The system transforms raw outage event data into structured outage records by changing parameters such as calculating outage duration from timestamps, determining user impact counts, and computing cost metrics. These parameter transformations automatically generate accurate outage information without manual intervention.
4Productivity
If automated monitoring is implemented to retrieve outage data from servers, then the productivity of outage tracking improves, but the device complexity increases
Solution Approach 1:
The monitored server automatically transmits outage event data to the monitoring application without requiring manual data collection. The server serves itself by generating and sending the outage information, which dramatically improves productivity while the monitoring application handles the automated processing and record generation.
Data Source
AI summary
Server outage data is automatically created and managed. Outage data is automatically retrieved from one or more servers at which an outage is detected by an agent installed on the server. The agent may search for outage event data and transmit the data to a monitoring application. The monitoring application receives the event data and creates an outage record from the data. Server contents, such as the number of users having account data on the server, can be determined either before or after the outage has occurred. Once the outage data and server contents are known, the cost and impact of the outage for each particular server can be determined. The cost of a server outage may be determined based on the outage record and the server data identifying resources of the server, such as user account data.


