Self-Healing Cluster Node Replacement for Availability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Cloud-based storage systems face performance impairments due to planned maintenance downtimes and other limitations, which existing automated tools are unable to effectively address, especially in software-defined storage platforms.

Innovation Solution

The implementation of self-healing techniques that include a processor-based virtual infrastructure monitoring entity detecting malfunctioning components in a cluster computing environment, selecting appropriate server types for replacements, creating and configuring new virtual infrastructure servers, and deploying replacement components while notifying the cluster monitoring entity for integration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If cloud-based infrastructure is used for software-defined storage systems, then scalability and flexibility are improved, but planned maintenance downtimes and limitations impair system performance and availability

Engineering Contradiction:
Improvescalability and flexibilityVSAvoidsystem availability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system performs preliminary actions by detecting malfunctioning components before they cause complete system failure. The automated mitigation process proactively replaces failed components with new virtual infrastructure servers, preventing service interruption and maintaining system availability while preserving the scalability of cloud-based infrastructure

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements self-service through automated detection and mitigation of component failures. The self-healing process automatically identifies malfunctioning components, provisions replacement servers, and restores service without human intervention, thereby maintaining high availability while leveraging the flexible cloud infrastructure

Inventive Principle:
Principle #25Self-service

2Reliability

If automated detection and mitigation processes are implemented, then system reliability and availability are improved, but system complexity increases

Engineering Contradiction:
Improvesystem availabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system merges multiple functions into integrated monitoring and mitigation processes. The cluster monitoring entity combines component failure detection, replacement server provisioning, and service restoration into a unified automated system, improving reliability while managing complexity through functional integration rather than separate discrete systems

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system introduces an intermediary automated mitigation process that sits between the cloud infrastructure and the storage services. This intermediary layer handles the complexity of failure detection and recovery, shielding the core storage services from complexity while maintaining high availability through automated intervention

Inventive Principle:
Principle #24Intermediary (Mediator)

3Device complexity

If manual intervention is used to address component failures, then system complexity is reduced, but downtime and productivity loss increase

Engineering Contradiction:
Improvemaintenance complexityVSAvoidsystem uptime
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The system enables self-service by automatically detecting component failures and executing mitigation processes without human intervention. The automated provisioning of replacement servers and restoration of services eliminates manual labor requirements, maintaining low operational complexity while maximizing system uptime and productivity through continuous automated operation

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements feedback mechanisms where the cluster monitoring entity continuously monitors component health and automatically triggers mitigation processes upon detecting failures. This closed-loop feedback system enables automated response to failures, reducing both manual complexity and downtime by instantly reacting to component malfunctions without waiting for human intervention

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12265445B2Detection and mitigation of malfunctioning components in a cluster computing environment
Publication Date: 2025.04.01 DELL PROD LP
  • US12265445B2 patent drawing
  • US12265445B2 patent drawing
  • US12265445B2 patent drawing

AI summary

Techniques are provided for detection and mitigation of malfunctioning components in a cluster computing environment. One method comprises obtaining, by a virtual infrastructure monitor, from a cluster monitor, an indication of a malfunctioning component in a cluster computing environment; selecting a virtual infrastructure server type for a replacement virtual infrastructure server based on a type of the malfunctioning component; creating a replacement virtual infrastructure server based on the selected virtual infrastructure server type and properties of a virtual infrastructure server associated with the malfunctioning component; applying settings to the replacement virtual infrastructure server according to rules for the replacement virtual infrastructure server; deploying a replacement component on the replacement virtual infrastructure server; and providing a notification to the cluster monitor of the replacement component and credentials of the replacement component. The cluster monitor may add the replacement component to the cluster computing environment responsive to the notification.