Self-Healing Cluster Node Replacement for Availability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Cloud-based storage systems face performance impairments due to planned maintenance downtimes and other limitations, which existing automated tools are unable to effectively address, especially in software-defined storage platforms.
Innovation Solution
The implementation of self-healing techniques that include a processor-based virtual infrastructure monitoring entity detecting malfunctioning components in a cluster computing environment, selecting appropriate server types for replacements, creating and configuring new virtual infrastructure servers, and deploying replacement components while notifying the cluster monitoring entity for integration.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If cloud-based infrastructure is used for software-defined storage systems, then scalability and flexibility are improved, but planned maintenance downtimes and limitations impair system performance and availability
Solution Approach 1:
The system performs preliminary actions by detecting malfunctioning components before they cause complete system failure. The automated mitigation process proactively replaces failed components with new virtual infrastructure servers, preventing service interruption and maintaining system availability while preserving the scalability of cloud-based infrastructure
Solution Approach 2:
The system implements self-service through automated detection and mitigation of component failures. The self-healing process automatically identifies malfunctioning components, provisions replacement servers, and restores service without human intervention, thereby maintaining high availability while leveraging the flexible cloud infrastructure
2Reliability
If automated detection and mitigation processes are implemented, then system reliability and availability are improved, but system complexity increases
Solution Approach 1:
The system merges multiple functions into integrated monitoring and mitigation processes. The cluster monitoring entity combines component failure detection, replacement server provisioning, and service restoration into a unified automated system, improving reliability while managing complexity through functional integration rather than separate discrete systems
Solution Approach 2:
The system introduces an intermediary automated mitigation process that sits between the cloud infrastructure and the storage services. This intermediary layer handles the complexity of failure detection and recovery, shielding the core storage services from complexity while maintaining high availability through automated intervention
3Device complexity
If manual intervention is used to address component failures, then system complexity is reduced, but downtime and productivity loss increase
Solution Approach 1:
The system enables self-service by automatically detecting component failures and executing mitigation processes without human intervention. The automated provisioning of replacement servers and restoration of services eliminates manual labor requirements, maintaining low operational complexity while maximizing system uptime and productivity through continuous automated operation
Solution Approach 2:
The system implements feedback mechanisms where the cluster monitoring entity continuously monitors component health and automatically triggers mitigation processes upon detecting failures. This closed-loop feedback system enables automated response to failures, reducing both manual complexity and downtime by instantly reacting to component malfunctions without waiting for human intervention
Data Source
AI summary
Techniques are provided for detection and mitigation of malfunctioning components in a cluster computing environment. One method comprises obtaining, by a virtual infrastructure monitor, from a cluster monitor, an indication of a malfunctioning component in a cluster computing environment; selecting a virtual infrastructure server type for a replacement virtual infrastructure server based on a type of the malfunctioning component; creating a replacement virtual infrastructure server based on the selected virtual infrastructure server type and properties of a virtual infrastructure server associated with the malfunctioning component; applying settings to the replacement virtual infrastructure server according to rules for the replacement virtual infrastructure server; deploying a replacement component on the replacement virtual infrastructure server; and providing a notification to the cluster monitor of the replacement component and credentials of the replacement component. The cluster monitor may add the replacement component to the cluster computing environment responsive to the notification.


