Case-Based Inference for Computing Facility State Management
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
As distributed computing systems grow in size and complexity, managing undesirable operational states becomes increasingly challenging due to the vast amount of event messages generated, which can lead to data loss and downtime, and existing monitoring tools often rely on manual intervention prone to errors and delays.
Innovation Solution
The implementation of case-based inference systems that maintain a database of previously handled operational states and remedial actions to facilitate automated or semi-automated diagnosis and correction of undesirable states, using similarity and desirability metrics to drive the system back to an acceptable operational state.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual monitoring and intervention methods are used to manage undesirable operational states, then system administrators can directly address operational issues, but errors and delays in detecting and addressing problems increase
Solution Approach 1:
The system performs self-diagnosis and self-management by automatically detecting undesirable operational states, retrieving relevant cases from the database, and applying appropriate remedial actions without requiring manual administrator intervention for every incident
Solution Approach 2:
The system continuously monitors operational states, compares them against stored case knowledge, and automatically adjusts system state based on feedback from similar historical cases, creating a closed-loop control mechanism that improves response accuracy and speed
2Productivity
If the size and complexity of distributed computing systems increase to provide enormous computational bandwidth, then system capacity and performance improve, but the difficulty of system management and the volume of event messages increase
Solution Approach 1:
The case-based management system serves multiple functions: it stores historical operational data, performs pattern recognition, retrieves relevant cases, and executes remedial actions, providing a universal management framework that scales with system complexity
Solution Approach 2:
The system creates and maintains a database of copied case records from historical operational states, allowing administrators to replicate successful remediation strategies across similar future incidents without re-analyzing problems from scratch
3Loss of time
If automated monitoring systems are implemented to reduce manual intervention, then response speed improves, but the complexity of the monitoring system increases
Solution Approach 1:
The system performs preliminary actions by pre-storing operational cases and their remedial solutions in a database during normal operation, so that when undesirable states occur, the system can quickly retrieve and apply pre-prepared solutions without complex real-time analysis
Solution Approach 2:
The case database serves as an intermediary between system monitoring and remedial actions, storing structured operational knowledge that mediates between detecting problems and applying solutions, simplifying the automation logic
Data Source
AI summary
The current document is directed to automatically, semi-automatically, and/or manually monitoring a computing facility to detect and address undesirable operational states in computing facilities, including large distributed computing systems. The currently disclosed monitoring methods and systems employ case-based inference to diagnose and ameliorate undesirable operational states. In disclosed implementations, a database is maintained to store and provide access to records of previously handled undesirable operational states and the actions taken to remediate the undesirable operational state is maintained in order to facilitate case-based reasoning inference.


