Complex Event Processing for Automated Fault Diagnosis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern computer systems face challenges in managing dynamic resources and handling changing user needs and system faults, leading to inefficiencies in fault rectification and disaster recovery, which are often labor-intensive and costly.
Innovation Solution
A complex event processing system that analyzes events in real-time, enabling self-management capabilities such as self-configuring, self-healing, and self-protecting through a network of agents and an analysis engine that monitors and adjusts resource behavior to minimize downtime and ensure continuous service.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional exception-catching schemes are used to handle faults, then fault-tolerant capability is improved, but system complexity and labor intensity increase
Solution Approach 1:
The system implements self-service through automated fault detection, diagnosis, and resolution mechanisms. The complex event processing system automatically monitors system state, identifies fault patterns, and executes remediation actions without human intervention, allowing the system to heal itself and reducing operational complexity despite enhanced fault tolerance
Solution Approach 2:
The system applies preliminary action by pre-defining fault patterns and their corresponding remediation policies before faults occur. The complex event processing system continuously compares real-time events against a library of known fault patterns, enabling rapid automatic response when patterns are detected, thus improving reliability while maintaining manageable system complexity through preparation
2Reliability
If extensive error handling capabilities are implemented, then system reliability is improved, but ease of operation deteriorates
Solution Approach 1:
The system automates the entire error handling process through self-service mechanisms. The complex event processing system continuously monitors events, automatically detects fault patterns, selects appropriate remediation policies, and executes corrective actions without requiring operator intervention, thereby maintaining high reliability while preserving ease of operation
Solution Approach 2:
The system implements continuous feedback loops where the complex event processing system monitors system state, compares it against known fault patterns, and automatically adjusts system behavior in response to detected anomalies. This closed-loop feedback mechanism ensures reliable error handling while keeping the system easy to operate by eliminating manual error management tasks
3Loss of time
If real-time event analysis is implemented, then fault diagnosis speed is improved, but device complexity increases
Solution Approach 1:
The system uses copying by maintaining a library of replicated fault patterns and their corresponding remediation policies. The complex event processing system compares real-time events against these pre-stored pattern copies, enabling rapid fault diagnosis without complex real-time analysis algorithms, thus reducing diagnosis time while managing system complexity through pattern replication
Solution Approach 2:
The system applies preliminary action by pre-processing and storing fault patterns in a reusable library before they are needed for diagnosis. The complex event processing system has these patterns prepared and indexed in advance, allowing rapid matching and diagnosis when faults occur, thereby reducing fault diagnosis time while keeping the processing system complexity manageable through upfront preparation
4Reliability
If automated self-healing mechanisms are deployed, then service continuity is improved, but system complexity increases
Solution Approach 1:
The system implements self-service through automated self-healing mechanisms where the complex event processing system autonomously detects service disruptions, identifies appropriate remediation policies from stored patterns, and executes corrective actions without human intervention. This enables continuous service operation while managing complexity through automation of healing processes
Solution Approach 2:
The system applies preliminary action by pre-defining and storing remediation policies for various fault patterns in an accessible library. When service disruptions occur, the complex event processing system quickly retrieves and executes the appropriate pre-prepared remediation actions, ensuring service continuity while keeping the automation system complexity manageable through advance preparation of recovery procedures
Data Source
AI summary
A computer system deploys monitoring agents that monitor the status and health of the computing resources. An analysis engine aggregates and analyzes event information from monitoring agents in order to support self-configuration, self-healing, self-optimization, and self-protection for managing the computer resources. If the analysis engine determines that a computing resource for a software application is approaching a critical status, the analysis engine may issue a command to that computing resource in accordance with a selected policy based on a detected event pattern. The command may indicate how the computing resource should change its behavior in order to minimize downtime for the software application as supported by that computing resource. The computer system may also support a distributed approach with a plurality of servers interacting with a central engine to manage the computer resources located at the servers.


