Microkernel Management System Automated Failure Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In datacenter and cloud computing environments, existing management systems face challenges in efficiently monitoring and automatically recovering from failures of their internal subsystems, leading to potential service disruptions and inefficiencies in resource management.
Innovation Solution
A management system is implemented with a microkernel architecture, service managers, and a meta-model that enables automated failure recovery, dynamic resource allocation, and service level management through a Service Oriented Architecture (SOA) framework, using a profile creator to specify deployment architectures and manage resources across different service levels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional management systems are used for each subset of the data center environment, then specific management functions can be performed, but the number of management systems proliferates and system complexity increases
Solution Approach 1:
The patent consolidates multiple specialized management systems into a single unified management system that can handle diverse management functions through a common architecture. This unified system uses a shared database, common user interface, and integrated service management framework to eliminate the proliferation of separate management systems while maintaining all necessary management capabilities.
Solution Approach 2:
The management system is designed with universal components that can perform multiple management functions. The system includes a framework that supports various service types (IT services, telecommunications services, etc.) and can manage different resources (hardware, software, network) through a single platform, making each component multi-functional rather than specialized.
2Device complexity
If manual monitoring and recovery processes are used, then system simplicity is maintained, but service availability decreases due to delayed failure recovery
Solution Approach 1:
The system performs preliminary actions by pre-configuring recovery procedures, service level agreements, and failure response protocols before failures occur. When failures are detected, pre-defined recovery actions are automatically executed, eliminating the need for manual intervention and ensuring rapid service restoration while maintaining system reliability.
Solution Approach 2:
The monitoring system implements continuous feedback loops that automatically detect service failures, assess their impact on service levels, and trigger appropriate recovery actions. The system monitors service performance metrics in real-time and uses this feedback to automatically initiate remediation processes, ensuring rapid response to failures without requiring complex manual monitoring procedures.
3Reliability
If automated failure recovery is implemented, then service availability improves, but the complexity of the management system increases
Solution Approach 1:
The management system is segmented into modular components with distinct responsibilities: service definition modules, monitoring modules, analysis modules, and remediation modules. Each component is independently designed and can be configured separately, allowing automated failure recovery functionality to be added without fundamentally complicating the overall system architecture. The modular design enables selective activation of automation features.
Solution Approach 2:
The system introduces an intermediary service level analysis component that sits between the monitoring system and the remediation system. This intermediary analyzes monitoring data, determines whether failures warrant automated recovery actions, and coordinates the execution of remediation procedures. This intermediary layer simplifies the automation logic by providing a decision-making buffer rather than requiring direct complex automation rules.
4Measurement precision
If extensive monitoring of internal subsystems is performed, then failure detection accuracy improves, but the complexity of monitoring increases
Solution Approach 1:
The system extracts and monitors only the critical service level indicators and key performance metrics that are essential for detecting failures affecting service delivery. Rather than monitoring all internal subsystem parameters, the system selectively monitors extracted metrics that directly correlate with service health, achieving high failure detection accuracy while keeping the monitoring mechanism simple and focused.
Data Source
AI summary
Systems and methods for automated failure recovery of subsystems of a management system are described. The subsystems are built and modeled as services, and their management, specifically their failure recovery, is done in a manner similar to that of services and resources managed by the management system. The management system consists of a microkernel, service managers, and management services. Each service, whether a managed service or a management service, is managed by a service manager. The service manager itself is a service and so is in turn managed by the microkernel. Both managed services and management services are monitored via in-band and out-of-band mechanisms, and the performance metrics and alerts are transported through an event system to the appropriate service manager. If a service fails, the service manager takes policy-based remedial steps including, for example, restarting the failed service.


