Service Model Flight Recorder for Event Replay
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current Service Impact Management (SIM) systems only allow visualization of the service model and impacting problems at a specific point in time, failing to provide insight into the sequence of events leading to service interruptions, which are often caused by cascades of multiple failures over time.
Innovation Solution
A method to record changes to a service impact model represented in a directed acyclic graph (DAG), enabling users to replay and visualize the chain of events leading to system outages, and allowing for retroactive calculation of metrics like MTTR and MTBF, with features like VCR-like playback and snapshot recording to diagnose business service disruptions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If Service Impact Management systems only visualize the service model at a specific point in time, then the system complexity is reduced and ease of operation is improved, but the ability to diagnose the sequence of events leading to service interruptions is lost
Solution Approach 1:
The system performs preliminary recording of service model states and events at multiple time points, storing this data in advance for later playback. This allows the system to capture the complete sequence of events leading to service interruptions without requiring complex real-time processing during diagnosis.
Solution Approach 2:
The system creates copies of the service model state at different time points and stores them as replayable recordings. These copies enable users to review historical states and event sequences without affecting the current operational system, thus maintaining simplicity while preserving diagnostic information.
2Measurement precision
If the system records and stores service model changes for replay, then the diagnostic capability is improved, but the device complexity increases
Solution Approach 1:
The service model is segmented into discrete states and events that can be independently recorded and replayed. This segmentation allows the system to manage complexity by breaking down the service model into manageable units while maintaining the ability to reconstruct complete event sequences for precise diagnostics.
3Measurement precision
If the system provides VCR-like playback functionality, then the ability to identify cause of business impacting events is improved, but the ease of operation deteriorates due to increased complexity
Solution Approach 1:
The system introduces a playback intermediary layer that separates the complex recording/replay functionality from the user interface. This intermediary handles the complexity of managing recorded states and events, while presenting a simplified VCR-like control interface to users, thus maintaining ease of operation while enabling precise event identification.
Data Source
AI summary
A method, system and medium for recording events in a system management environment is described. As system events are detected in an enterprise computing environment they are stored in a manner allowing them to be “replayed” either forward or reverse to assist a system administrator or other user to determine the chain of events that affected the enterprise. The system engineer and business process owner are therefore presented with pertinent information for monitoring, administrating and diagnosing system activities and their correlation to business services.


