Automated Virtual Machine Replacement Orchestration in Cloud Clusters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Cloud-based virtual machines experience non-recoverable failures due to hardware issues, leading to data vulnerability and increased support cases, resulting in degraded user experiences and high costs.
Innovation Solution
An orchestration engine automates the replacement of failed virtual machines by triggering data protection and rebalancing jobs, ensuring data integrity and minimizing downtime through automated processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If virtual machines are deployed in cloud-based environments, then resource flexibility and scalability are improved, but reliability deteriorates due to higher failure rates
Solution Approach 1:
The system performs preliminary actions by automatically detecting virtual machine failures and initiating replacement processes before complete system degradation occurs. The orchestration engine proactively monitors health status and triggers replacement workflows to prevent total service disruption.
Solution Approach 2:
The system implements self-service through automated failure detection and replacement orchestration. The orchestration engine autonomously manages the entire replacement process including resource allocation, data migration, and service restoration without requiring manual intervention, enabling the system to self-heal from failures.
2Ease of operation
If manual replacement processes are used for failed virtual machines, then control and monitoring are improved, but loss of time and productivity deteriorate
Solution Approach 1:
The orchestration engine performs self-service by automatically detecting failures, determining replacement needs, and executing the entire replacement workflow without human intervention. This eliminates manual operations while maintaining full control through automated monitoring and decision-making.
Solution Approach 2:
The system ensures continuity of useful action by maintaining ongoing monitoring of virtual machine health and continuously ready-to-execute replacement processes. When failures occur, the replacement workflow is already prepared and can be initiated immediately, minimizing disruption to business operations.
3Productivity
If automated replacement processes are implemented, then productivity and reliability are improved, but device complexity increases
Solution Approach 1:
The orchestration engine provides multi-functionality by handling multiple operations including failure detection, health status monitoring, replacement initiation, resource allocation, and service validation through a single unified system. This consolidates what would otherwise require multiple separate systems into one manageable platform.
4Ease of operation
If virtual machine failures are not quickly addressed, then operational simplicity is maintained, but loss of information and data vulnerability increase
Solution Approach 1:
The system performs preliminary actions by continuously monitoring virtual machine health status and being ready to execute replacement workflows immediately upon detecting failures. This proactive approach prevents data vulnerability by ensuring rapid response before failures can propagate or cause information loss.
Data Source
AI summary
The technology described herein is directed towards automating the replacement of a virtual machine when the hardware underlying the virtual machine fails, including in a cloud computing environment in which nodes in a cluster map to virtual machines being deployed within that cloud provider. An automated workflow to perform cluster self-healing is started upon detection of an unrecoverable instance failure of a virtual machine, e.g., because of underlying hardware failure. The failed virtual machine is terminated, and a new, replacement virtual machine that matches characteristics of the failed virtual machine is created to join the cluster. Data of the failed node is re-protected, such as by restoring data maintained with a protection scheme to remaining virtual machines of the cluster. When the data is re-protected and the replacement virtual machine has joined the cluster, the data is rebalanced across the cluster nodes, including to the new virtual machine.


